model releasexAI

xAI releases Grok 4.20 with 2M context window and native reasoning capabilities

TL;DR

xAI released Grok 4.20 on March 31, 2026, its flagship model featuring a 2 million token context window, $2 per million input tokens and $6 per million output tokens pricing, and toggleable reasoning capabilities. The model includes web search functionality at $5 per 1,000 queries and claims industry-leading speed with low hallucination rates.

2 min read
0

Grok 4.20 — Quick Specs

Context window2000K tokens
Input$2/1M tokens
Output$6/1M tokens

xAI Releases Grok 4.20 With 2M Context and Reasoning Mode

xAI released Grok 4.20 on March 31, 2026. The model features a 2 million token context window, priced at $2 per million input tokens and $6 per million output tokens. Web search functionality costs $5 per 1,000 queries. Knowledge cutoff is September 1, 2025.

Key Specifications

Context and Pricing:

  • Context window: 2,000,000 tokens
  • Input pricing: $2/M tokens
  • Output pricing: $6/M tokens
  • Web search: $5 per 1,000 queries

Capabilities: Grok 4.20 includes native agentic tool-calling abilities and an optional reasoning mode that can be enabled or disabled via API parameter. When enabled, the model exposes its step-by-step reasoning process through a reasoning_details array in the response. Users can preserve reasoning context across multi-turn conversations by passing the complete reasoning details back in subsequent requests.

Performance Claims

xAI claims Grok 4.20 delivers "industry-leading speed" combined with low hallucination rates and "strict prompt adherence." The company positions the model as producing "consistently precise and truthful responses," though specific benchmark scores against competing models have not been disclosed.

Technical Details

The model is accessible through OpenRouter, which normalizes API requests and responses across multiple providers and includes fallback routing to maximize uptime. OpenRouter's documentation indicates support for reasoning-enabled models with access to internal step-by-step thinking before final outputs.

API integration follows standard patterns with optional OpenRouter-specific headers for leaderboard attribution. Third-party SDK support is available through OpenRouter's framework documentation.

What This Means

Grok 4.20 enters a competitive market where context window size has become a commodity feature—Claude 3.5 Sonnet and other models already offer 200K context, and some specialty models exceed 1M. The 2M window is a notable advantage but not unprecedented. The model's value proposition rests on claimed speed advantages and agentic capabilities rather than raw context size alone. The optional reasoning mode, similar to features in newer OpenAI and Anthropic models, allows developers to choose between faster inference and detailed reasoning transparency based on use case requirements. Pricing at $2/$6 positions it at the premium end of the market—significantly higher than open-source alternatives but within range of other flagship models. The lack of disclosed benchmark scores leaves performance claims unverified; independent evaluation will be necessary to validate xAI's claims about hallucination rates and prompt adherence relative to competitors.

Related Articles

model release

H Company Releases Holo4 Agentic Models, Scoring 61.7% on OSWorld 2.0 at a Fraction of Frontier Cost

H Company has released Holo4, a series of agentic models in 27B dense and 35B-A3B Mixture of Experts sizes that operate across GUIs, code, MCP, and APIs using a single interface. The models score 61.7% (27B) and 30.9% (35B-A3B) on OSWorld 2.0, trailing closed frontier models like Opus 5.5 (81.8%) but at far lower cost.

model release

Nvidia Releases Nemotron 3 Diarization, a Free 100M-Parameter Model That Tracks 8 Speakers in Real Time

Nvidia released Nemotron 3 Diarization, a free 100-million-parameter model that identifies who is speaking in real time across up to eight participants. It leads the VoiceArena Diarization Benchmark v1 with a 14.7% error rate, cutting errors by 41% versus its predecessor.

model release

Apple Releases LensVLM-9B, a 9B Vision-Language Model That Selectively Decompresses Text Images

Apple has released LensVLM-9B, a 9-billion-parameter vision-language model fine-tuned from Qwen3.5-9B-Base that processes documents as compressed images, selectively expanding only relevant pages to full resolution. The model supports 5x, 10x, and 15x compression ratios and is available under Apple's Machine Learning Research Model License.

model release

Meta Releases Muse Glimmer 30B, an Open-Weight Agentic Model for Consumer Hardware

Meta Superintelligence Labs has released Muse Glimmer 30B, a dense open-weight model distilled from its larger Muse Spark system and tuned for agentic workflows on consumer hardware. The model supports 131K context, image understanding, and over 100 languages at $0.30/$1.10 per 1M input/output tokens.

Comments

Loading...