OpenAI cut the API price of its budget GPT-5.6 Luna model by 80% on July 30, just three weeks after the model launched at its original rate on July 9. Input tokens dropped from $1.00 to $0.20 per million, and output tokens fell from $6.00 to $1.20 per million. The mid-tier GPT-5.6 Terra model also got a 20% cut, from $2.50/$15.00 to $2.00/$12.00 per million input and output tokens, while the flagship GPT-5.6 Sol tier held steady at $5.00 and $30.00.
The cuts land Luna roughly five times cheaper than Anthropic's Claude Haiku 4.5 on input costs and about four times cheaper on output, according to a comparison of published rate cards. Terra now undercuts Claude Sonnet 5's standard pricing of $3 input and $15 output per million tokens, set to take effect August 31.
Anthropic's answer wasn't a price cut — it was more model for the same price
Related: Google's Cheap New Gemini Model Targets Crypto's Trading Bots
Anthropic took a different approach days earlier, launching Claude Opus 5 at the same combined rate as its predecessor, Opus 4.8 — $5 input and $25 output per million tokens — rather than cutting the sticker price. The effect is similar to a price cut in practice: buyers get a materially more capable model at an unchanged rate, lowering the effective cost per unit of output quality rather than the quoted price itself. OpenAI's Sol flagship remains more expensive than Opus 5 on output pricing even after the Luna and Terra adjustments.
Beyond headline rates, OpenAI is also leaning on volume-based discounts to win high-throughput customers: batch and flex processing on Luna now costs $0.10 input and $0.60 output per million tokens, a 50% discount, while cached input tokens carry a 90% discount, bringing Luna's cheapest effective rate down to $0.02 per million tokens. Industry trackers estimate that prices across leading U.S. models are down almost 25% collectively since mid-July.
The squeeze is showing up in enterprise budgets too
The price war is unfolding alongside a separate shift toward spend discipline: OpenAI introduced dashboard-enforced monthly spending limits on July 22, and Amazon has simultaneously capped its own internal AI spending. Together, the moves suggest the market for large language model inference is starting to behave like a genuine commodity market — one where providers compete on cost-per-token for routine workloads like classification, extraction and multi-turn agent tasks, even as they continue to charge a premium for frontier-capability models. For an industry that has increasingly overlapped with crypto through the same GPU and data-center infrastructure that Bitcoin miners have been pivoting toward, falling inference prices are a direct read on how much pricing power AI compute providers still have left.