model releasexAI

xAI Releases Grok 4.6, a 1.5T-Parameter Model Powering New 'Grok Bot' AI Teammate Product

TL;DR

xAI released Grok 4.6, a confirmed 1.5T-parameter model built on Grok 4.5 with heavier training on long-horizon agentic tasks. It powers the newly launched Grok Bot product and scores 61 on Artificial Analysis's Intelligence Index at $2/$6 per 1M input/output tokens — well below frontier competitors.

3 min read
0

xAI shipped Grok 4.6 on August 12, 2026, a confirmed 1.5-trillion-parameter model that the company says builds on Grok 4.5 with heavier emphasis on long-running agents and more ambitious interactive and visual work. The release powers Grok Bot, a new early-beta AI teammate product that signs into users' tools, performs tasks autonomously, and returns with finished work.

Benchmarks and Pricing

Independent evaluation from Artificial Analysis places Grok 4.6 at 61 on its Intelligence Index — roughly matching GPT-5.6 Sol Max and trailing Claude Opus/Fable. The model posted 88.4% on Terminal-Bench v2.1 and a 1753 Elo on GDPval-AA v2. On Artificial Analysis's proprietary AA-Briefcase benchmark, which tests long-horizon agentic knowledge work on a private, contamination-resistant test set, Grok 4.6 made what AA described as large gains while costing substantially less than competing frontier models.

Pricing is $2 per 1M input tokens and $6 per 1M output tokens, according to Artificial Analysis — materially cheaper than rival frontier models at similar capability levels. Early users on Code Arena placed it near GPT-5.6 Sol and Claude Fable on web development tasks, and Cognition has already made it available inside Devin.

Training Disclosure

xAI's technical notes, reproduced by AI News, describe a longer supplemental training run than Grok 4.5 used, incorporating curated model-generated data for reasoning and technical concepts, higher-quality engineering data, and an improved optimizer and training recipe. Grok 4.5 was then used to regenerate SFT trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, with model-based filtering to remove problematic traces. The resulting checkpoint underwent agentic RL training across knowledge work, general coding, and domain-specific environments spanning kernel optimization, web development, and computer-aided design.

xAI also reported that Grok 4.6 exhibits more self-testing behavior during long-running tasks compared to its predecessor. Elon Musk stated that Grok 4.7 is already in training, with initial training complete and supplemental training planned on SpaceX internal data — a detail that underscores the tightening relationship between xAI and SpaceX referenced in coverage of the release.

Competitive Context

Grok 4.6 arrived alongside a cluster of other frontier releases the same week, including Alibaba's Qwen3.8-Max (a 2.4T total / 95B active parameter open-weight MoE model, text-only in its initial drop), DeepSeek's V4 Pro GA (priced at roughly $0.435 per 1M input tokens and $0.87 per 1M output tokens), and Microsoft's first from-scratch reasoning model, MAI-Thinking-1, now in Foundry. Both Cognition and Musk himself have acknowledged Grok 4.6 as arguably the second-best knowledge work model currently available, behind Claude's latest, though xAI is positioning it as the leader on price-to-performance.

What This Means

Grok 4.6's real significance may not be the Intelligence Index score — which lands mid-pack among frontier models — but the pairing of competitive agentic benchmarks with aggressive pricing and a dedicated "AI teammate" product wrapper. The AI-teammate category, following Claude Tag and Block's Buzz, has been searching for a clear leader; Grok Bot's early reception suggests xAI may have found a workable formula by combining a cost-efficient model with tool-using autonomy. The disclosed training process — reusing Grok 4.5 to regenerate SFT data and layering agentic RL across specialized environments — also signals that further gains in this category will likely come from environment design and self-distillation rather than raw parameter scaling, a trend echoed in DeepSeek's V4 Pro results this same week.

Related Articles

Comments

Loading...