OpenAI announced on Thursday drastic price reductions across the GPT-5.6 family just weeks after launch. Luna, the cheapest model in the line, dropped 80%: the cost per million input tokens fell from $1 to $0.20, and output from $6 to $1.20. Terra, the mid-tier model, saw a 20% cut, while Sol gained a Fast mode delivering up to 2.5 times more throughput at double the rate, aimed at latency-sensitive workloads. The new rates already apply automatically to Codex and to usage accounting within ChatGPT Work.
The decision did not come out of nowhere. It arrives at a moment when the model market has shifted its focus from raw capability to cost per unit of intelligence. Open-weight models, many produced in China, now offer competitive performance for a fraction of the price — and pressure on more expensive proprietary APIs has become unavoidable. The cut is also a response to enterprise pushback: companies that began questioning the real cost of running agents and automation at scale pushed OpenAI to lower the barrier for intensive use.
A technical detail makes the announcement even more interesting: part of the reduction was funded by efficiency gains that GPT-5.6 Sol itself achieved by autonomously rewriting OpenAI's inference infrastructure. In other words, the model helped optimize the system that serves it — and the savings were passed on to customers as lower prices. This inverts an old industry logic, where quality improvements almost always came with higher prices.
The price war is not new, but the pace is changing. Before, cuts happened when a competitor released something better; now, they happen when the cost gap grows too large to ignore. If an open-weight model delivers 80% of the performance at 10% of the price, proprietary pricing must justify itself another way — through reliability or exclusive ecosystem features. OpenAI has clearly chosen to compete on cost and integrated experience.
The question that remains is how far this decline can go. Token prices still have room to shrink, but at some point the marginal cost of inference becomes the floor. And when the cheapest model becomes cheap enough to use without thinking twice, the entire market changes nature — what is contested is no longer price, but the utility of what is built on top. Perhaps that is the real revolution today's cut signals.
For developers building products on top of APIs, the price cut is more than welcome news: it is a change in the economic equation of what is viable to build. Applications that previously relied on aggressive caching or careful modeling to reduce the number of calls can now afford to be more generous, and agents running in long loops suddenly become cheaper to operate. The move also signals the direction the largest AI companies are taking to defend against open-weight competition: instead of trying to sell exclusivity at a premium price, they bet on volume and integration, delivering the low cost that the market has already come to consider normal. The question is whether this race to the bottom on prices will leave enough margin to fund the next generation of research.
Sources: VentureBeat, ExplainX AI, ZeroHedge
✓ Independent sources cross-checked and verified before publishing