xAI Releases Grok 4.7 at $2/$6 per Million Tokens, Trails Claude and GPT-6 on Benchmarks
xAI has launched Grok 4.7 at $2 per million input tokens and $6 per million output tokens, undercutting Western rivals on price. But independent benchmarks show it trailing Claude Fable 5.1 and GPT-6 by a wide margin, especially in agentic coding.
xAI has launched Grok 4.7, its newest model for coding and knowledge work, at $2 per million input tokens and $6 per million output tokens. That pricing undercuts most Western frontier models and sits closer to rates typically seen from Chinese labs — a positioning that independent benchmarks suggest may reflect the model's actual competitive tier rather than aggressive discounting.
Benchmark performance
On the Artificial Analysis Intelligence Index (v4.3.2), which aggregates ten benchmarks into a single score, Grok 4.7 posts a score of 46, landing mid-pack. Claude Fable 5.1 and GPT-6 both lead the index with scores of 53. According to Artificial Analysis, Grok 4.7's two highest reasoning levels perform roughly the same on this index, suggesting limited gains from pushing the model to its most expensive reasoning setting.
The gap widens sharply in agentic coding. On Terminal-Bench 4.0, a benchmark measuring autonomous coding task completion, Grok 4.7 scores just 26 percent. GPT-6 Astra hits 60 percent and Claude Fable 5.1 reaches 55 percent — both more than double Grok 4.7's result. Even DeepSeek V4.1 Flash, a lower-cost Chinese model, edges past Grok 4.7 with 27 percent on the same test.
What xAI claims
According to xAI, Grok 4.7 is built on a larger base model than its predecessor, trained with longer reinforcement learning runs, and designed to better verify its own output before returning a response. The company positions it as its most capable model yet for coding and knowledge work, though the independent benchmark data does not support a frontier-level claim relative to Claude Fable 5.1 or GPT-6.
Grok 4.7 is available now through the Grok API, in Cursor, and through Grok Build. xAI has not disclosed the model's context window size or parameter count.
What this means
Grok 4.7's benchmark results put xAI in an awkward spot. The pricing — $2/$6 per million tokens — is aggressive enough to compete with budget-focused Chinese models like DeepSeek, but the coding performance doesn't back up a premium positioning against Claude Fable 5.1 or GPT-6. A 26 percent score on Terminal-Bench 4.0, worse than DeepSeek V4.1 Flash's 27 percent, is a weak result for a model xAI is marketing specifically for coding work.
The low price likely reflects this reality rather than a strategic discount: xAI appears to be pricing Grok 4.7 in line with its measured capability rather than its ambition. For developers choosing a model for agentic coding tasks, the benchmark gap — GPT-6 Astra and Claude Fable 5.1 both scoring more than double Grok 4.7's Terminal-Bench result — is difficult to ignore regardless of the price advantage. Grok 4.7 may still find a market among cost-sensitive users running simpler tasks, but it does not close the gap with the current frontier on complex reasoning or autonomous coding work.
Related Articles
xAI Ships Grok 4.7, Cuts Price to $1.60/$4.80 per 1M Tokens With 500K Context
xAI has released Grok 4.7, the successor to Grok 4.6, listed on OpenRouter with a 500K token context window and pricing of $1.60 per 1M input tokens and $4.80 per 1M output tokens. The company claims improvements in long-running software engineering tasks, self-verification, and professional document drafting.
OpenAI's GPT-6 Astra Beats Pokémon in 18 Hours, Scores 62.7% on ARC-AGI-3
GPT-6 Astra completed Pokémon FireRed in 18 hours 12 minutes, five times faster than its predecessor, and scored 62.7% on ARC-AGI-3 versus 7.78% for GPT-5.6 Sol. The model also ran a 141-hour Minecraft session and finished Fallout 3 in roughly 59 hours, according to independent testers.
Qwen3.8-Omni-Flash Prices Multimodal AI at $0.15/$0.47 per Million Tokens, Undercutting Gemini Flash by 5x
Alibaba's Qwen team released Qwen3.8-Omni-Flash, a multimodal model for AI agents that processes audio and video with a 1 million token context window. Pricing undercuts Google's Gemini 3.8 Flash by roughly 5x on input and 8x on output, according to Qwen.
Cactus Compute Releases Needle 3, a Sub-30MB On-Device Model for Tool Calling and Extraction
Cactus Compute has released Needle 3, a foundation model compressed to 8-29 MB that runs entirely on-device for tool calling, structured data extraction, and text embedding. The company claims it beats models 10x its size on mobile tool calls while running on hardware as small as microcontrollers.
Comments
Loading...