changelogOpenAI

OpenAI Slashes GPT-5.6 Luna Pricing by 80%, Cuts Terra by 20%

TL;DR

OpenAI cut GPT-5.6 Luna pricing by 80 percent to $0.20 per million input tokens and $1.20 per million output tokens, while Terra dropped 20 percent to $2/$12. The company attributes the cuts to infrastructure efficiency gains and mounting price competition, particularly from Chinese providers.

3 min read
1

OpenAI Slashes GPT-5.6 Luna Pricing by 80 Percent

OpenAI cut pricing on its smallest GPT-5.6 model, Luna, by 80 percent effective July 30. The new rate is $0.20 per million input tokens and $1.20 per million output tokens, down from previous pricing that OpenAI did not disclose in its announcement.

GPT-5.6 Terra, the mid-tier model, gets a smaller 20 percent cut, now priced at $2 per million input tokens and $12 per million output tokens. GPT-5.6 Sol, the top-tier model in the lineup, keeps its existing pricing unchanged.

All three models remain available through ChatGPT Work, Codex, and the OpenAI API.

OpenAI's performance claims

According to OpenAI, Luna now matches the performance of what the company calls "leading models from a year ago." OpenAI claims a task that would have cost $1 to run on those older models now costs roughly 6 cents on Luna — a claimed cost reduction of more than 90 percent — while running nearly nine times faster. OpenAI has not published independent benchmark scores to substantiate these comparisons, so the figures should be treated as company claims rather than verified third-party results.

Why OpenAI says it can afford the cuts

OpenAI attributes the price reductions largely to infrastructure gains driven by GPT-5.6 Sol. The company says Sol optimized its own GPU software, cutting deployment costs by 20 percent, and improved token generation throughput by more than 15 percent using speculative decoding — a technique where a smaller draft model predicts tokens that a larger model then verifies, reducing compute overhead per output token. OpenAI has not detailed how much of Sol's optimization work was automated versus human-supervised.

Market context

The cuts land amid intensifying price competition in the AI model market. Chinese labs — including DeepSeek, Alibaba's Qwen, Moonshot AI, and Zhipu AI — have pushed aggressive low-cost pricing on models with competitive benchmark performance, pressuring Western providers to respond. Microsoft has also begun promoting its own MAI models as cheaper alternatives to OpenAI's lineup, despite OpenAI and Microsoft's close commercial partnership.

What this means

This is a pricing and efficiency story, not a capability release — OpenAI is not claiming Luna is a new or smarter model, only that it now costs dramatically less to run at a stated performance level comparable to older frontier systems. The 80 percent cut to Luna signals OpenAI's intent to compete directly on cost in high-volume, latency-sensitive use cases (chatbots, agents, batch processing) where Chinese providers have gained ground on price-to-performance.

The deeper risk sits with OpenAI's own financials. The company has committed to massive infrastructure spending tied to expected revenue growth from frontier model pricing. An industry-wide price war — even one driven by genuine efficiency gains like speculative decoding and self-optimized GPU software — compresses margins across the board. If rivals match these cuts, the sustainability question shifts from "can OpenAI afford to charge less" to "can any lab recover its infrastructure costs at these price points." Watch whether Anthropic, Google DeepMind, and xAI respond with matching cuts to their comparable tiers in the coming weeks.

Related Articles

product update

OpenAI Launches Agents API in Public Beta, Exposing Codex Infrastructure to Developers

OpenAI has released the Agents API in public beta, giving developers access to the same cloud infrastructure that powers Codex and ChatGPT. The API supports long-running agents, parallel tool use, and sub-agent delegation, with billing based solely on token usage.

benchmark

AWS Benchmark: OpenAI's GPT-5.6 Luna Beats GPT-5.4 Mini on Cost-Per-Correct-Answer Despite Similar List Price

An AWS blog post using an open-source benchmarking harness finds that GPT-5.6 Luna, Terra, and Sol on Amazon Bedrock deliver lower cost-per-correct-answer than OpenAI's cost-optimized GPT-5.4 Mini and Nano, once accuracy, token efficiency, and agent turn counts are factored in. The analysis also cites a July 30, 2026 price cut of up to 80% for GPT-5.6 Luna on Amazon Bedrock.

product update

OpenAI Launches ChatGPT for Financial Services to Automate Wall Street Analyst Work

OpenAI launched ChatGPT for Financial Services, a tailored enterprise product built with design partners Morgan Stanley and Evercore that automates research, financial analysis, and pitchbook creation. The tool, powered by GPT-6 Astra, targets tasks traditionally performed by Wall Street's junior analysts and associates.

model release

OpenAI Launches GPT-Image-2.5 in Two Variants, Cuts Latency Up to 50%

OpenAI has released GPT-Image-2.5 in two variants — Flare and Sunburst — promising faster generation, more precise multi-step editing, and new 'xhigh' and 'max' quality tiers. Both models use identical token pricing but land at the top of Arena's preliminary text-to-image leaderboard.

Comments

Loading...