OpenAI has cut the price developers pay for its most capable model by more than 20%. The reduction runs for three months, and it arrives while the company is fighting on two fronts: Anthropic above it on capability claims, and a growing set of cheap Chinese models underneath.
The new numbers
GPT-5.6 Sol now costs $4 per million input tokens and $20 per million output tokens for standard short-context work. On Thursday, the same job cost $5 and $30.
The cut applies to the API, and it's rolling out across eligible plans for credits on ChatGPT Work, OpenAI's agentic product, and Codex, its coding tool. If you pay for Pro, Plus or Business, nothing about your bill changes. This one is for the people whose costs scale with how much they actually use.
A pattern, not a one-off
Late last month OpenAI trimmed its mid-tier Terra model by 20% and took 80% off Luna, the budget option. Cutting the cheap end first and then the flagship looks less like a promotion and more like a company reworking its entire price ladder.
The competitive math is worth spelling out. Anthropic lists Claude Fable 5 at $10 in and $50 out, and Claude Opus 5 at $5 and $25. At $4 and $20, Sol is now the cheapest frontier tier on the board. Whether that wins any business is a separate question — teams rarely switch models over a line in a pricing table, because the cost of re-testing and re-tuning everything usually swamps the savings.
Where the savings actually show up
Output tokens are where the money goes in agent and coding work. The model spends most of its effort generating, not reading, so dropping the output rate from $30 to $20 changes the arithmetic on exactly the workloads that have been too expensive to run at full scale.
That's the practical read here. A long multi-step agent task that cost $30 to run now costs $20. Multiply that across a team running hundreds of those a day and the difference stops being a rounding error.
Read the expiry date
Three months. That's the part worth writing down somewhere. This is promotional pricing, and any budget built on $4 and $20 needs an answer for what happens when the window closes. If you've been sitting on a deployment because the token bill looked ugly, you now have a defined stretch of time to find out whether the work pays for itself at the lower rate — and whether it still does at the old one.
What this says about the market
For most of the last two years, AI labs competed on capability and let price trail behind. Cuts like this suggest the top models have converged enough that cost per token is doing real competitive work.
That matters beyond OpenAI's margins. Cheaper frontier access means a small team can build on the best available model instead of settling for a mid-tier one, which is the kind of change that shows up later in what actually gets built.
The thing to watch: whether $4 and $20 quietly becomes the list price in November, or whether the old numbers come back.
The qualifier worth reading twice
These rates cover standard short-context use. That phrase matters. Long-context requests are priced on their own terms, so if your workload involves stuffing huge documents or long agent histories into every call, don't assume a clean 20% off your invoice. Check your own usage against the table.
It's also worth separating the two halves of this announcement, because they aren't the same thing. The API cut is immediate and mechanical — the unit price drops. The ChatGPT Work and Codex piece arrives as credits across eligible plans, which is value applied against what you already pay rather than a lower rate. Same press release, two very different lines on a finance spreadsheet.
Image: Alicia Christin Gerald, via Pexels





