xAI Releases Grok 4.6, a 1.5T-Parameter Model Powering New 'Grok Bot' AI Teammate Product
xAI released Grok 4.6, a confirmed 1.5T-parameter model built on Grok 4.5 with heavier training on long-horizon agentic tasks. It powers the newly launched Grok Bot product and scores 61 on Artificial Analysis's Intelligence Index at $2/$6 per 1M input/output tokens — well below frontier competitors.
xAI shipped Grok 4.6 on August 12, 2026, a confirmed 1.5-trillion-parameter model that the company says builds on Grok 4.5 with heavier emphasis on long-running agents and more ambitious interactive and visual work. The release powers Grok Bot, a new early-beta AI teammate product that signs into users' tools, performs tasks autonomously, and returns with finished work.
Benchmarks and Pricing
Independent evaluation from Artificial Analysis places Grok 4.6 at 61 on its Intelligence Index — roughly matching GPT-5.6 Sol Max and trailing Claude Opus/Fable. The model posted 88.4% on Terminal-Bench v2.1 and a 1753 Elo on GDPval-AA v2. On Artificial Analysis's proprietary AA-Briefcase benchmark, which tests long-horizon agentic knowledge work on a private, contamination-resistant test set, Grok 4.6 made what AA described as large gains while costing substantially less than competing frontier models.
Pricing is $2 per 1M input tokens and $6 per 1M output tokens, according to Artificial Analysis — materially cheaper than rival frontier models at similar capability levels. Early users on Code Arena placed it near GPT-5.6 Sol and Claude Fable on web development tasks, and Cognition has already made it available inside Devin.
Training Disclosure
xAI's technical notes, reproduced by AI News, describe a longer supplemental training run than Grok 4.5 used, incorporating curated model-generated data for reasoning and technical concepts, higher-quality engineering data, and an improved optimizer and training recipe. Grok 4.5 was then used to regenerate SFT trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, with model-based filtering to remove problematic traces. The resulting checkpoint underwent agentic RL training across knowledge work, general coding, and domain-specific environments spanning kernel optimization, web development, and computer-aided design.
xAI also reported that Grok 4.6 exhibits more self-testing behavior during long-running tasks compared to its predecessor. Elon Musk stated that Grok 4.7 is already in training, with initial training complete and supplemental training planned on SpaceX internal data — a detail that underscores the tightening relationship between xAI and SpaceX referenced in coverage of the release.
Competitive Context
Grok 4.6 arrived alongside a cluster of other frontier releases the same week, including Alibaba's Qwen3.8-Max (a 2.4T total / 95B active parameter open-weight MoE model, text-only in its initial drop), DeepSeek's V4 Pro GA (priced at roughly $0.435 per 1M input tokens and $0.87 per 1M output tokens), and Microsoft's first from-scratch reasoning model, MAI-Thinking-1, now in Foundry. Both Cognition and Musk himself have acknowledged Grok 4.6 as arguably the second-best knowledge work model currently available, behind Claude's latest, though xAI is positioning it as the leader on price-to-performance.
What This Means
Grok 4.6's real significance may not be the Intelligence Index score — which lands mid-pack among frontier models — but the pairing of competitive agentic benchmarks with aggressive pricing and a dedicated "AI teammate" product wrapper. The AI-teammate category, following Claude Tag and Block's Buzz, has been searching for a clear leader; Grok Bot's early reception suggests xAI may have found a workable formula by combining a cost-efficient model with tool-using autonomy. The disclosed training process — reusing Grok 4.5 to regenerate SFT data and layering agentic RL across specialized environments — also signals that further gains in this category will likely come from environment design and self-distillation rather than raw parameter scaling, a trend echoed in DeepSeek's V4 Pro results this same week.
Related Articles
xAI's Grok 4.6 Matches Claude and GPT-5.6 on Benchmarks, Costs 60% Less
xAI's Grok 4.6 ties OpenAI's GPT-5.6 Sol on the Artificial Analysis Intelligence Index with a score of 61, trailing only Anthropic's Claude Opus 5 and Claude Fable 5. Pricing remains at $2/$6 per million tokens, undercutting both competitors by more than 60 percent.
xAI Launches Grok 4.6, Claims Parity with GPT-5.6 Sol and Near-Parity with Claude Fable 5
xAI released Grok 4.6, claiming intelligence on par with OpenAI's GPT-5.6 Sol and just one point behind Anthropic's Claude Fable 5 Max on the Artificial Analysis Intelligence Index. The model is priced at $2 per million input tokens and $6 per million output tokens, with a faster variant at double that rate.
xAI Launches Grok 4.6 With 500K Token Context Window
xAI has released Grok 4.6, a text-and-image model featuring a 500K token context window. The model is priced at $2.00 per million input tokens and $6.00 per million output tokens, and is available now through OpenRouter's API.
DeepSeek Releases V4 Pro 0813 With 1.05M Token Context Window, Priced at $0.43/M Input
DeepSeek has shipped the general availability release of DeepSeek V4 Pro, codenamed 0813, featuring a 1,049,000-token context window. The mixture-of-experts model is priced at $0.43 per million input tokens and $0.87 per million output tokens, and is live now on OpenRouter.
Comments
Loading...