model release

xAI's Grok 4.6 Matches Claude and GPT-5.6 on Benchmarks, Costs 60% Less

TL;DR

xAI's Grok 4.6 ties OpenAI's GPT-5.6 Sol on the Artificial Analysis Intelligence Index with a score of 61, trailing only Anthropic's Claude Opus 5 and Claude Fable 5. Pricing remains at $2/$6 per million tokens, undercutting both competitors by more than 60 percent.

2 min read
0

Grok 4.6 Ties GPT-5.6 Sol on Intelligence Index

xAI released Grok 4.6, a model that scores 61 points on the Artificial Analysis Intelligence Index — matching OpenAI's GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5 (63) and Claude Fable 5 (62). The score marks a five-point jump over its predecessor, Grok 4.5, and puts Grok 4.6 in the top tier of frontier models as measured by the index, which aggregates multiple benchmarks into a single composite score.

The bigger story is price. Grok 4.6 is priced at $2 per million input tokens and $6 per million output tokens — more than 60 percent cheaper than Claude Opus 5 ($5/$25 per million tokens) and GPT-5.6 Sol ($5/$30 per million tokens). xAI has kept pricing unchanged from the previous Grok generation despite the performance gain.

Strong Agentic Performance

Grok 4.6's biggest gains show up in agentic evaluation. On GDPval-AA v2, a benchmark designed to measure real-world knowledge work performed on a computer, Grok 4.6 ranks second overall with an Elo score of 1,753, behind only Claude Opus 5. According to Artificial Analysis, Grok 4.6 completes complex multi-step tasks in an average of 53 steps, compared to roughly 103 steps for Claude Opus 5 — suggesting notably higher efficiency per task, though step count alone doesn't capture task quality or accuracy.

Availability

Grok 4.6 is live now via xAI's API, Cursor, Grok Build, and third-party platforms including OpenRouter, Vercel, and Cloudflare. For the first week following release, xAI is offering double usage quotas for Grok Build and Cursor users, according to the company.

Context window size and training cutoff date have not been disclosed by xAI as of publication.

What This Means

Grok 4.6 closes the performance gap with Anthropic and OpenAI's top models while significantly undercutting them on cost — a combination that matters most for teams running high-volume agentic workloads, where token spend adds up quickly across many-step tasks. The efficiency gap on GDPval-AA v2 (53 steps versus Claude Opus 5's 103) is notable, but Elo-based step counts don't necessarily translate to equivalent output quality, and independent verification of xAI's claims beyond the Artificial Analysis benchmark is still limited.

For developers choosing between frontier models, Grok 4.6 now presents a credible price-performance alternative rather than a clear performance leader. Whether it holds up outside controlled benchmarks — particularly on reliability, tool use, and long-context tasks — will determine if the price advantage translates into real adoption away from Claude and GPT-5.6 Sol.

Related Articles

model release

xAI Launches Grok 4.6, Claims Parity with GPT-5.6 Sol and Near-Parity with Claude Fable 5

xAI released Grok 4.6, claiming intelligence on par with OpenAI's GPT-5.6 Sol and just one point behind Anthropic's Claude Fable 5 Max on the Artificial Analysis Intelligence Index. The model is priced at $2 per million input tokens and $6 per million output tokens, with a faster variant at double that rate.

model release

xAI Launches Grok 4.6 With 500K Token Context Window

xAI has released Grok 4.6, a text-and-image model featuring a 500K token context window. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens, and is available now through OpenRouter's API.

model release

Nvidia Releases Nemotron 3.5 Lightning: A 31.6B-Parameter Open Model Built for Speed, Not Peak Intelligence

Nvidia's Nemotron 3.5 Lightning, a 31.6B-parameter open-weight model with only 3.6B active parameters, matches OpenAI's gpt-oss-120b on the Artificial Analysis Intelligence Index while delivering the fastest throughput in its class at nearly 670 tokens per second. The model posts especially large gains on agentic benchmarks, beating both gpt-oss-120b and the larger Nemotron 3 Super.

model release

Meta Releases Muse Glimmer, a 30B Open-Weight Agent Model That Runs on a Single RTX 3090

Meta released Muse Glimmer, a 30B-parameter open-weight model under Apache 2.0 built for always-on local agents, alongside a promise to release Muse Spark 1.2 weights soon. The model runs on a single RTX 3090 and scores 35 on Artificial Analysis's Intelligence Index.

Comments

Loading...