model releasexAI

xAI Releases Grok 4.6, a 1.5T-Parameter Model Powering New 'Grok Bot' AI Teammate Product

TL;DR

xAI released Grok 4.6, a confirmed 1.5T-parameter model built on Grok 4.5 with heavier training on long-horizon agentic tasks. It powers the newly launched Grok Bot product and scores 61 on Artificial Analysis's Intelligence Index at $2/$6 per 1M input/output tokens — well below frontier competitors.

3 min read
1

xAI shipped Grok 4.6 on August 12, 2026, a confirmed 1.5-trillion-parameter model that the company says builds on Grok 4.5 with heavier emphasis on long-running agents and more ambitious interactive and visual work. The release powers Grok Bot, a new early-beta AI teammate product that signs into users' tools, performs tasks autonomously, and returns with finished work.

Benchmarks and Pricing

Independent evaluation from Artificial Analysis places Grok 4.6 at 61 on its Intelligence Index — roughly matching GPT-5.6 Sol Max and trailing Claude Opus/Fable. The model posted 88.4% on Terminal-Bench v2.1 and a 1753 Elo on GDPval-AA v2. On Artificial Analysis's proprietary AA-Briefcase benchmark, which tests long-horizon agentic knowledge work on a private, contamination-resistant test set, Grok 4.6 made what AA described as large gains while costing substantially less than competing frontier models.

Pricing is $2 per 1M input tokens and $6 per 1M output tokens, according to Artificial Analysis — materially cheaper than rival frontier models at similar capability levels. Early users on Code Arena placed it near GPT-5.6 Sol and Claude Fable on web development tasks, and Cognition has already made it available inside Devin.

Training Disclosure

xAI's technical notes, reproduced by AI News, describe a longer supplemental training run than Grok 4.5 used, incorporating curated model-generated data for reasoning and technical concepts, higher-quality engineering data, and an improved optimizer and training recipe. Grok 4.5 was then used to regenerate SFT trajectories across reasoning efforts, agent harnesses, and domains including STEM, software engineering, and knowledge work, with model-based filtering to remove problematic traces. The resulting checkpoint underwent agentic RL training across knowledge work, general coding, and domain-specific environments spanning kernel optimization, web development, and computer-aided design.

xAI also reported that Grok 4.6 exhibits more self-testing behavior during long-running tasks compared to its predecessor. Elon Musk stated that Grok 4.7 is already in training, with initial training complete and supplemental training planned on SpaceX internal data — a detail that underscores the tightening relationship between xAI and SpaceX referenced in coverage of the release.

Competitive Context

Grok 4.6 arrived alongside a cluster of other frontier releases the same week, including Alibaba's Qwen3.8-Max (a 2.4T total / 95B active parameter open-weight MoE model, text-only in its initial drop), DeepSeek's V4 Pro GA (priced at roughly $0.435 per 1M input tokens and $0.87 per 1M output tokens), and Microsoft's first from-scratch reasoning model, MAI-Thinking-1, now in Foundry. Both Cognition and Musk himself have acknowledged Grok 4.6 as arguably the second-best knowledge work model currently available, behind Claude's latest, though xAI is positioning it as the leader on price-to-performance.

What This Means

Grok 4.6's real significance may not be the Intelligence Index score — which lands mid-pack among frontier models — but the pairing of competitive agentic benchmarks with aggressive pricing and a dedicated "AI teammate" product wrapper. The AI-teammate category, following Claude Tag and Block's Buzz, has been searching for a clear leader; Grok Bot's early reception suggests xAI may have found a workable formula by combining a cost-efficient model with tool-using autonomy. The disclosed training process — reusing Grok 4.5 to regenerate SFT data and layering agentic RL across specialized environments — also signals that further gains in this category will likely come from environment design and self-distillation rather than raw parameter scaling, a trend echoed in DeepSeek's V4 Pro results this same week.

Related Articles

model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

model release

Meta Releases Muse Glimmer 30B, an Open-Weight Agentic Model for Consumer Hardware

Meta Superintelligence Labs has released Muse Glimmer 30B, a dense open-weight model distilled from its larger Muse Spark system and tuned for agentic workflows on consumer hardware. The model supports 131K context, image understanding, and over 100 languages at $0.30/$1.10 per 1M input/output tokens.

model release

AionLabs Launches Aion 3.5, a Multi-Model Storytelling System Built on GLM

AionLabs has released Aion 3.5, a collaborative multi-model system for roleplaying and storytelling built on the GLM model family. It offers a 262K token context window at $3 per 1M input tokens and $6 per 1M output tokens.

model release

Upstage Releases Solar Mini 4: 35B MoE Model with 524K Context at $0.05/$0.20 per Million Tokens

Upstage has released Solar Mini 4, a compact mixture-of-experts model with 35B total parameters, 3B active parameters, and a 524K token context window. The model targets agentic workloads and is priced at $0.05 per 1M input tokens and $0.20 per 1M output tokens, a promotional 50% discount off standard rates.

Comments

Loading...