xAI's Grok 4.6 Matches Claude and GPT-5.6 on Benchmarks, Costs 60% Less
xAI's Grok 4.6 ties OpenAI's GPT-5.6 Sol on the Artificial Analysis Intelligence Index with a score of 61, trailing only Anthropic's Claude Opus 5 and Claude Fable 5. Pricing remains at $2/$6 per million tokens, undercutting both competitors by more than 60 percent.
Grok 4.6 Ties GPT-5.6 Sol on Intelligence Index
xAI released Grok 4.6, a model that scores 61 points on the Artificial Analysis Intelligence Index — matching OpenAI's GPT-5.6 Sol and trailing only Anthropic's Claude Opus 5 (63) and Claude Fable 5 (62). The score marks a five-point jump over its predecessor, Grok 4.5, and puts Grok 4.6 in the top tier of frontier models as measured by the index, which aggregates multiple benchmarks into a single composite score.
The bigger story is price. Grok 4.6 is priced at $2 per million input tokens and $6 per million output tokens — more than 60 percent cheaper than Claude Opus 5 ($5/$25 per million tokens) and GPT-5.6 Sol ($5/$30 per million tokens). xAI has kept pricing unchanged from the previous Grok generation despite the performance gain.
Strong Agentic Performance
Grok 4.6's biggest gains show up in agentic evaluation. On GDPval-AA v2, a benchmark designed to measure real-world knowledge work performed on a computer, Grok 4.6 ranks second overall with an Elo score of 1,753, behind only Claude Opus 5. According to Artificial Analysis, Grok 4.6 completes complex multi-step tasks in an average of 53 steps, compared to roughly 103 steps for Claude Opus 5 — suggesting notably higher efficiency per task, though step count alone doesn't capture task quality or accuracy.
Availability
Grok 4.6 is live now via xAI's API, Cursor, Grok Build, and third-party platforms including OpenRouter, Vercel, and Cloudflare. For the first week following release, xAI is offering double usage quotas for Grok Build and Cursor users, according to the company.
Context window size and training cutoff date have not been disclosed by xAI as of publication.
What This Means
Grok 4.6 closes the performance gap with Anthropic and OpenAI's top models while significantly undercutting them on cost — a combination that matters most for teams running high-volume agentic workloads, where token spend adds up quickly across many-step tasks. The efficiency gap on GDPval-AA v2 (53 steps versus Claude Opus 5's 103) is notable, but Elo-based step counts don't necessarily translate to equivalent output quality, and independent verification of xAI's claims beyond the Artificial Analysis benchmark is still limited.
For developers choosing between frontier models, Grok 4.6 now presents a credible price-performance alternative rather than a clear performance leader. Whether it holds up outside controlled benchmarks — particularly on reliability, tool use, and long-context tasks — will determine if the price advantage translates into real adoption away from Claude and GPT-5.6 Sol.
Related Articles
OpenAI's GPT-6 Astra Reportedly Automates AI Engineering Tasks at Under $6 an Hour, According to Latent Space Testing
A Latent Space report describes GPT-6 Astra, a new OpenAI model the blog says can autonomously handle AI engineering tasks—training models, labeling data, deploying systems—at an estimated cost of under $6 per hour. The claims, including 97.6% on FrontierMath and 99.9% on ARC-AGI-3, come from independent blog testing rather than an official OpenAI announcement.
Meta Releases Muse Spark 1.3, a Free Multimodal Reasoning Model with 1M-Token Context
Meta has released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window, listed as free on OpenRouter. The model targets long-running agentic, multi-agent, and coding workflows, though audio input support remains incomplete.
OpenAI Rates Upcoming Astra Model 'Critical' Risk for Cyber Capabilities — Its Highest Tier Ever
OpenAI says its unreleased Astra model is the first to trigger a 'critical' cybersecurity rating under its Preparedness Framework, capable of finding and chaining unknown vulnerabilities without human guidance. The company calls it simultaneously its most dangerous and safest model, while a new architecture detail raises questions about how well its reasoning can still be monitored.
OpenAI's GPT-6 Astra Cuts Hallucinations, But Indirect Prompt Injection Attacks Still Succeed 8.5% of the Time
OpenAI's new GPT-6 Astra model shows major improvements in hallucination rates and jailbreak resistance over predecessor GPT-5.6 Sol, according to OpenAI's system card. However, indirect prompt injection attacks hidden in documents still succeed 8.5% of the time in external testing by Gray Swan, down from 27% but still above rival Claude Opus 5's 4.8% rate.
Comments
Loading...