IBM Releases Granite 4.2 8B, a Dense Reasoning Model with 131K Context and Three Thinking Modes
IBM has released Granite 4.2 8B, a dense reasoning model built for math, code generation, and agentic workflows. The model supports 131K context, 12 languages, and three switchable reasoning modes, priced at $0.10 per 1M input tokens and $0.15 per 1M output tokens.
Granite 4.2 8B — Quick Specs
IBM has released Granite 4.2 8B, a dense reasoning model designed for mathematics, code generation, multilingual dialogue, and agentic workflows requiring multi-step reasoning. The model is available now through OpenRouter and IBM's model distribution channels.
Specifications
Granite 4.2 8B ships with a 131,000-token context window and is priced at $0.10 per 1 million input tokens and $0.15 per 1 million output tokens. As the name indicates, the model has 8 billion parameters and uses a dense architecture rather than a mixture-of-experts design.
The model supports three inference modes: full reasoning, low-effort reasoning, and non-thinking. This gives developers control over the latency-versus-reasoning-depth tradeoff on a per-request basis, using a single set of model weights rather than requiring separate model variants.
Granite 4.2 8B supports 12 languages: English, German, Spanish, French, Japanese, Portuguese, Arabic, Czech, Italian, Korean, Dutch, and Chinese. IBM positions the model for agentic workflows, code generation, and mathematical reasoning tasks alongside standard multilingual dialogue.
No independent benchmark scores were disclosed in IBM's release materials or the OpenRouter listing. Training cutoff date has not been disclosed.
Availability
Model weights are available for download, and the model can be tested through OpenRouter's playground and API. Uptime and latency data across providers is still limited, with OpenRouter noting fewer than three days of tracked performance data at time of listing.
What this means
Granite 4.2 8B continues IBM's strategy of shipping small, dense, cost-efficient models aimed at enterprise deployment rather than chasing frontier-scale parameter counts. At $0.10/$0.15 per 1M tokens, it undercuts most mid-size reasoning models on price while offering a 131K context window that is competitive with larger, more expensive alternatives.
The three-mode inference switch (full, low-effort, non-thinking) is the most practically interesting feature here. It lets teams dial reasoning depth up or down without swapping models, which matters for cost control in production agentic systems where not every step needs maximum deliberation.
The absence of published benchmark scores makes it difficult to place Granite 4.2 8B against comparable dense models like Llama 3.1 8B or Qwen2.5 7B on standard evaluations. Until third-party benchmarks emerge, claims about its reasoning and coding capability should be treated as unverified. Its value proposition for now rests on price, context length, and multilingual coverage rather than demonstrated leaderboard performance.
Related Articles
IBM Releases Granite 4.2 Open-Weight Models With Agentic RL Training and 512K Context
IBM has released Granite 4.2, a family of open-weight language models in 3B, 8B, and 30B parameter sizes, trained on roughly 15 trillion tokens with context windows up to 512,000 tokens. The 8B and 30B variants underwent additional 'agentic RL' training for tool use, code execution, and web search.
IBM Releases Granite 4.2, Its First Reasoning-Focused LLM Family in 3B, 8B, and 30B Sizes
IBM has published a technical breakdown of Granite 4.2, its first dense, decoder-only reasoning model family, released in 3B, 8B, and 30B sizes. The models are pre-trained on roughly 15 trillion tokens, extended to a 512K-token context window, and post-trained with a multi-stage RL pipeline that includes agentic tool-use training for the 8B and 30B variants.
Alibaba Releases Qwen3.8 Flash, a Multimodal Reasoning Model with 1M-Token Context
Alibaba has released Qwen3.8 Flash, a multimodal reasoning model with a 1 million token context window, aimed at coding, agentic workflows, and visual/document analysis. It's priced at $0.16 per 1M input tokens and $0.47 per 1M output tokens through Alibaba Cloud International.
Alibaba Releases Qwen3.8-Flash-Next: 125B-Parameter MoE Model Matches Larger Rivals at $0.16/$0.47 per Million Tokens
Alibaba's Qwen team released Qwen3.8-Flash-Next, a 125-billion-parameter mixture-of-experts model that activates just 6 billion parameters per token and previews architecture planned for Qwen4. The model outperforms the much larger Qwen3.7-Plus at roughly one-ninth the training cost and ships at $0.16 per million input tokens and $0.47 per million output tokens.
Comments
Loading...