changelogOpenAI

OpenAI Slashes GPT-5.6 Luna Pricing by 80%, Cuts Terra by 20%

TL;DR

OpenAI cut GPT-5.6 Luna pricing by 80 percent to $0.20 per million input tokens and $1.20 per million output tokens, while Terra dropped 20 percent to $2/$12. The company attributes the cuts to infrastructure efficiency gains and mounting price competition, particularly from Chinese providers.

3 min read
1

OpenAI Slashes GPT-5.6 Luna Pricing by 80 Percent

OpenAI cut pricing on its smallest GPT-5.6 model, Luna, by 80 percent effective July 30. The new rate is $0.20 per million input tokens and $1.20 per million output tokens, down from previous pricing that OpenAI did not disclose in its announcement.

GPT-5.6 Terra, the mid-tier model, gets a smaller 20 percent cut, now priced at $2 per million input tokens and $12 per million output tokens. GPT-5.6 Sol, the top-tier model in the lineup, keeps its existing pricing unchanged.

All three models remain available through ChatGPT Work, Codex, and the OpenAI API.

OpenAI's performance claims

According to OpenAI, Luna now matches the performance of what the company calls "leading models from a year ago." OpenAI claims a task that would have cost $1 to run on those older models now costs roughly 6 cents on Luna — a claimed cost reduction of more than 90 percent — while running nearly nine times faster. OpenAI has not published independent benchmark scores to substantiate these comparisons, so the figures should be treated as company claims rather than verified third-party results.

Why OpenAI says it can afford the cuts

OpenAI attributes the price reductions largely to infrastructure gains driven by GPT-5.6 Sol. The company says Sol optimized its own GPU software, cutting deployment costs by 20 percent, and improved token generation throughput by more than 15 percent using speculative decoding — a technique where a smaller draft model predicts tokens that a larger model then verifies, reducing compute overhead per output token. OpenAI has not detailed how much of Sol's optimization work was automated versus human-supervised.

Market context

The cuts land amid intensifying price competition in the AI model market. Chinese labs — including DeepSeek, Alibaba's Qwen, Moonshot AI, and Zhipu AI — have pushed aggressive low-cost pricing on models with competitive benchmark performance, pressuring Western providers to respond. Microsoft has also begun promoting its own MAI models as cheaper alternatives to OpenAI's lineup, despite OpenAI and Microsoft's close commercial partnership.

What this means

This is a pricing and efficiency story, not a capability release — OpenAI is not claiming Luna is a new or smarter model, only that it now costs dramatically less to run at a stated performance level comparable to older frontier systems. The 80 percent cut to Luna signals OpenAI's intent to compete directly on cost in high-volume, latency-sensitive use cases (chatbots, agents, batch processing) where Chinese providers have gained ground on price-to-performance.

The deeper risk sits with OpenAI's own financials. The company has committed to massive infrastructure spending tied to expected revenue growth from frontier model pricing. An industry-wide price war — even one driven by genuine efficiency gains like speculative decoding and self-optimized GPU software — compresses margins across the board. If rivals match these cuts, the sustainability question shifts from "can OpenAI afford to charge less" to "can any lab recover its infrastructure costs at these price points." Watch whether Anthropic, Google DeepMind, and xAI respond with matching cuts to their comparable tiers in the coming weeks.

Comments

Loading...