model release

xAI's Grok 4.6 Matches Claude and GPT-5.6 on Benchmarks, Costs 60% Less

TL;DR

xAI's Grok 4.6 ties OpenAI's GPT-5.6 Sol on the Artificial Analysis Intelligence Index with a score of 61, trailing only Anthropic's Claude Opus 5 and Claude Fable 5. Pricing remains at $2/$6 per million tokens, undercutting both competitors by more than 60 percent.

2 min read
0

Grok 4.6 Ties GPT-5.6 Sol on Intelligence Index

xAI released Grok 4.6, a model that scores 61 points on the Artificial Analysis Intelligence Index — matching OpenAI's GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5 (63) and Claude Fable 5 (62). The score marks a five-point jump over its predecessor, Grok 4.5, and puts Grok 4.6 in the top tier of frontier models as measured by the index, which aggregates multiple benchmarks into a single composite score.

The bigger story is price. Grok 4.6 is priced at $2 per million input tokens and $6 per million output tokens — more than 60 percent cheaper than Claude Opus 5 ($5/$25 per million tokens) and GPT-5.6 Sol ($5/$30 per million tokens). xAI has kept pricing unchanged from the previous Grok generation despite the performance gain.

Strong Agentic Performance

Grok 4.6's biggest gains show up in agentic evaluation. On GDPval-AA v2, a benchmark designed to measure real-world knowledge work performed on a computer, Grok 4.6 ranks second overall with an Elo score of 1,753, behind only Claude Opus 5. According to Artificial Analysis, Grok 4.6 completes complex multi-step tasks in an average of 53 steps, compared to roughly 103 steps for Claude Opus 5 — suggesting notably higher efficiency per task, though step count alone doesn't capture task quality or accuracy.

Availability

Grok 4.6 is live now via xAI's API, Cursor, Grok Build, and third-party platforms including OpenRouter, Vercel, and Cloudflare. For the first week following release, xAI is offering double usage quotas for Grok Build and Cursor users, according to the company.

Context window size and training cutoff date have not been disclosed by xAI as of publication.

What This Means

Grok 4.6 closes the performance gap with Anthropic and OpenAI's top models while significantly undercutting them on cost — a combination that matters most for teams running high-volume agentic workloads, where token spend adds up quickly across many-step tasks. The efficiency gap on GDPval-AA v2 (53 steps versus Claude Opus 5's 103) is notable, but Elo-based step counts don't necessarily translate to equivalent output quality, and independent verification of xAI's claims beyond the Artificial Analysis benchmark is still limited.

For developers choosing between frontier models, Grok 4.6 now presents a credible price-performance alternative rather than a clear performance leader. Whether it holds up outside controlled benchmarks — particularly on reliability, tool use, and long-context tasks — will determine if the price advantage translates into real adoption away from Claude and GPT-5.6 Sol.

Related Articles

model release

Meta Releases Muse Glimmer 30B, an Open-Weight Agentic Model for Consumer Hardware

Meta Superintelligence Labs has released Muse Glimmer 30B, a dense open-weight model distilled from its larger Muse Spark system and tuned for agentic workflows on consumer hardware. The model supports 131K context, image understanding, and over 100 languages at $0.30/$1.10 per 1M input/output tokens.

model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

model release

Anthropic Launches Claude Opus 5.5 at 20% Lower List Price, Claims Parity with Claude Fable 5.1

Anthropic released Claude Opus 5.5, the first model in its new 5.5 family, cutting list pricing 20% to $4/$20 per 1M input/output tokens while claiming performance on par with Claude Fable 5.1. Independent analysis shows the cost savings largely disappear at maximum reasoning effort due to higher token consumption.

model release

Anthropic Ships Claude Opus 5.5, OpenAI Counters with GPT-6 Sol and Luna Hours Later, Triggering Sharp Price Cuts

Anthropic released Claude Opus 5.5 with a 20% price cut, and roughly an hour later OpenAI shipped GPT-6 Sol and GPT-6 Luna at roughly half the price of their GPT-5.6 predecessors. The releases follow Grok 4.7 and MiMo v2.6 from the previous day, intensifying competition among frontier model providers.

Comments

Loading...