Frontier AI prices halved again. Your AI bill probably won't.
Anthropic's Claude Opus 5.5 beats last quarter's top model at 60% less, and OpenAI cut GPT-6 Sol and Luna prices in half on the same day. Here is why cheaper tokens rarely mean a smaller bill, and how to spend the savings.
On September 22 Anthropic and OpenAI both released new models, and both led with price.
Anthropic’s Claude Opus 5.5 costs $4 per million input tokens and $20 per million output tokens, down from $5 and $25 for Opus 5. On Anthropic’s own figures it also beats Fable 5.1, the company’s most capable model, on agentic coding, scoring 66.4% against 55.8% on Terminal-Bench 4.0. Per token it is 60% cheaper than Fable 5.1. Cached input fell from $0.50 to $0.20 per million tokens, which matters more than the headline for anyone running long agent sessions.
OpenAI’s GPT-6 Sol and Luna went further on price. Sol now costs $2 and $10, half of what GPT-5.6 Sol cost. Luna, the high-volume tier meant for summarising and extraction, costs $0.10 and $0.50. The same day OpenAI opened broad access to GPT-6 Astra, its new top model, at $10 and $50.
Last quarter’s best model gets cheap fast
In June, Fable 5 launched as Anthropic’s top model at $10 and $50. In July, Opus 5 arrived close to Fable 5’s performance at half the price. In September, Opus 5.5 passed Fable 5.1 on several benchmarks at 60% less. Three times in one quarter, capability that cost a premium at launch reappeared one tier down, at a fraction of the price, within weeks or months.
Open models push from below. Moonshot’s Kimi K3 costs $3 and $15 through its own API, and anyone can host the weights. Every closed lab now sets prices with a free alternative a few places below it on the leaderboards.
Prices at the very top have not moved. Astra launched at the same $10 and $50 as Fable 5, and OpenAI says it meets its Critical threshold for cybersecurity capability, with the most sensitive defensive uses routed through a separate access programme. The newest top model still costs $50 per million output tokens and comes with access conditions, as it did in June.
Why your bill may not fall
Lower prices per token do not guarantee a lower bill, because the work companies give AI is changing faster than the price. A chatbot answering a question uses a few thousand tokens. An agent refactoring a codebase or working through a case file can use millions in a single session, and the more reliable models get, the longer the tasks people trust them with. A lot of companies will see token prices halve and AI spending double in the same year.
The discounts may not last. The Bank for International Settlements warned in June that the AI buildout is financed well ahead of revenue, and part of what makes these prices possible is investors’ willingness to fund that gap. Part of every discount is paid for by investors, and they can stop paying.
How to spend the savings well
Measure cost per completed task, not cost per token. OpenAI now uses this framing itself, claiming Sol finishes a task on its automation benchmark for $0.27. A cheaper model that needs three attempts costs more than a pricier one that gets it right the first time. The only numbers that count are the ones from your own workload.
Route by difficulty. Most business AI work is routine: classification, extraction, summaries, first drafts. That work belongs on the cheapest tier that does it acceptably, which as of this week means Luna-class prices. Save the expensive models for tasks where the better answer is worth the extra cost.
Use caching. Most agent and document work sends the same instructions and reference material with every request. On Opus 5.5, cached input costs $0.20 per million tokens against $4 for fresh input, a 95% discount on the repeated part. If your developers have not set up prompt caching, that is likely the largest saving available to you this month, larger than switching models.
Re-test every quarter. A model choice made in June is probably wrong by September. Keep a small set of real tasks from your business and rerun it when new models ship. It takes an afternoon, and one switch to a cheaper model that performs as well covers the time many times over.
And keep contracts short. Volume commitments negotiated at June prices look poor in September. If a vendor wants a long commitment, ask for pricing that follows their own list price down.

