model release

50 articles tagged with model release

October 6, 2026
model release

Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October

Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.

October 1, 2026
model releaseUnbiased

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

September 30, 2026
model release

Google Releases Gemini 4 Argon to Cybersecurity Partners, Claims Wins Over GPT-6 Astra

Google has released Gemini 4 Argon, its next-generation flagship AI model, to a small group of cybersecurity partners as part of a phased rollout. The company claims the model outperforms OpenAI's GPT-6 Astra on several coding and knowledge-work benchmarks, though full specifications remain undisclosed.

model release

Google DeepMind Releases Gemini 4 Argon, Expands Output Limit to 1M Tokens

Google DeepMind has released Gemini 4 Argon, a frontier model built for long-horizon reasoning with an industry-leading 1 million output token limit. The model is rolling out first to trusted cyber defenders through Google's Fairwind Program, with pricing set at $2 per million input tokens and $10 per million output tokens.

model release

Google Announces Gemini 4 Argon, Its New Frontier Model With 1M Output Tokens

Google has announced Gemini 4 Argon as its new frontier model, featuring a 1M output token limit (up from 64K) and claimed leads on coding, cybersecurity, and automation benchmarks. The model is rolling out first to Google AI Ultra subscribers and paid API customers.

September 29, 2026
model releaseOpenAI

OpenAI Releases GPT-6.1 Sol: Mid-Tier Model with 1.1M Context at $2/$10 per Million Tokens

OpenAI has released GPT-6.1 Sol, an incremental upgrade to GPT-6 Sol positioned below flagship GPT-6 Astra in its GPT-6 lineup. The model features a 1.1M token context window, priced at $2 per 1M input tokens and $10 per 1M output tokens, with claimed improvements in factual accuracy and instruction-following on agentic tasks.

September 28, 2026
model releaseOpenAI

OpenAI Scraps Release of GPT-6.1 Astra Over Safety Concerns

OpenAI confirmed it will not release GPT-6.1 Astra after the model failed to meet internal safety and alignment standards. The decision follows renewed industry-wide calls, including from Anthropic, to slow the pace of frontier model development.

model releaseAnthropic

Anthropic's Claude Sonnet 5.5 Launches on Amazon Bedrock and Claude Platform on AWS

Anthropic's Claude Sonnet 5.5 is now available on Amazon Bedrock and Claude Platform on AWS, positioned as a faster, lower-cost model for well-scoped coding and document tasks. It pairs with the recently released Claude Opus 5.5, which handles higher-judgment work.

model releaseAnthropic

Anthropic Releases Claude Sonnet 5.5, Nearly Matching Opus 5.5 at Up to 30% Lower Cost Per Task

Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family, delivering major coding and knowledge-work gains that approach flagship Opus 5.5 performance while using up to 30% fewer tokens per task. Per-token pricing stays unchanged at $2/$10 per million input/output tokens.

model releaseAnthropic

Anthropic Launches Claude Sonnet 5.5, Claims 30% Faster Performance at Lower Cost Than Predecessor

Anthropic has released Sonnet 5.5, the latest version of its mid-tier Claude model, claiming 30% faster performance and significantly lower token costs than its predecessor. The company says the model now outperforms Opus 5.5 on agentic coding tasks and carries cyber capabilities comparable to Opus 5.

September 23, 2026
model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

model release

AionLabs Launches Aion 3.5, a Multi-Model Storytelling System Built on GLM

AionLabs has released Aion 3.5, a collaborative multi-model system for roleplaying and storytelling built on the GLM model family. It offers a 262K token context window at $3 per 1M input tokens and $6 per 1M output tokens.

model releaseUpstage

Upstage Releases Solar Mini 4: 35B MoE Model with 524K Context at $0.05/$0.20 per Million Tokens

Upstage has released Solar Mini 4, a compact mixture-of-experts model with 35B total parameters, 3B active parameters, and a 524K token context window. The model targets agentic workloads and is priced at $0.05 per 1M input tokens and $0.20 per 1M output tokens, a promotional 50% discount off standard rates.

September 22, 2026
analysis

OpenRouter Listings Surface for Three Unannounced OpenAI Models: GPT-6 Sol Pro, Luna, and Luna Pro

OpenRouter's model directory listed three new entries—GPT-6 Sol Pro, GPT-6 Luna, and GPT-6 Luna Pro—attributed to OpenAI, but no pricing, context window, benchmark data, or official confirmation from OpenAI has surfaced.

model releaseAnthropic

Anthropic Ships Claude Opus 5.5, OpenAI Launches GPT-6 Sol and Luna — All Cheaper Than Predecessors

Anthropic released Claude Opus 5.5 at $4/$20 per million input/output tokens, undercutting Opus 5's $5/$25 pricing while claiming better agentic coding scores. OpenAI countered with GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50), both up to 50% cheaper than GPT-5.6's promotional rates.

model releaseOpenAI

OpenAI Launches GPT-6 Luna: Fast, Low-Cost Model With 1.1M Context Window

OpenAI has released GPT-6 Luna, the fast and cost-efficient entry in its new GPT-6 model family, featuring a 1.1M token context window and pricing starting at $0.10 per 1M input tokens. The model is positioned below GPT-6 Sol and GPT-6 Astra in OpenAI's tiered lineup.

model releaseOpenAI

OpenAI Launches GPT-6 Sol and GPT-6 Luna, Cutting API Prices 50% Versus GPT-5.6

OpenAI has released GPT-6 Sol and GPT-6 Luna, two new models that cost 50% less than their GPT-5.6 equivalents while claiming improved coding and computer-use performance. The models roll out today to ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users.

model releaseOpenAI

OpenAI's GPT-6 Sol and GPT-6 Luna Launch on Amazon Bedrock, Priced Below GPT-5.6

OpenAI has launched GPT-6 Sol and GPT-6 Luna on Amazon Bedrock, positioned below flagship GPT-6 Astra for recurring coding tasks and high-volume document processing respectively. Both models cost less per API call than their GPT-5.6 predecessors, though exact pricing figures were not disclosed.

model releaseAnthropic

Claude Opus 5.5 Launches on Amazon Bedrock, Anthropic's First Model in New 5.5 Family

Claude Opus 5.5, the first model in Anthropic's new Claude 5.5 family, is now live on Amazon Bedrock and Claude Platform on AWS. Anthropic claims the model does more with fewer tokens than Claude Opus 5, lowering average cost per task despite unchanged headline pricing tiers.

model releaseAnthropic

Anthropic Releases Claude Opus 5.5, Cuts Costs 40% While Matching Rival Fable 5.1

Anthropic has released Claude Opus 5.5, claiming performance parity with Claude Fable 5.1 at roughly 40% lower total operating cost than Opus 5. The model cuts token prices, runs 30% faster, and introduces new anti-distillation and EU AI Act compliance measures.

September 21, 2026
analysis

Xiaomi Lists Three New MiMo-V2.6 Models on OpenRouter: Pro, Flash, and Pro-UltraSpeed

Xiaomi has added three new entries to its MiMo model family on OpenRouter: MiMo-V2.6-Pro, MiMo-V2.6-Flash, and MiMo-V2.6-Pro-UltraSpeed. Full specifications, pricing, and benchmark results have not yet been disclosed.

model releaseXiaomi

Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output

Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.

model releaseXiaomi

Xiaomi Launches MiMo-V2.6-Pro, a 1T+ Parameter Model with 1M-Token Context

Xiaomi has released MiMo-V2.6-Pro, a flagship foundation model exceeding 1 trillion parameters with a 1M-token context window and native multimodal support. The model is priced at $0.435 per 1M input tokens and $0.87 per 1M output tokens, targeting agentic and long-horizon tasks.

model releasexAI

xAI Ships Grok 4.7, Cuts Price to $1.60/$4.80 per 1M Tokens With 500K Context

xAI has released Grok 4.7, the successor to Grok 4.6, listed on OpenRouter with a 500K token context window and pricing of $1.60 per 1M input tokens and $4.80 per 1M output tokens. The company claims improvements in long-running software engineering tasks, self-verification, and professional document drafting.

September 17, 2026
model release+1

Unbiased Launches Pareto, a $2.50/$7.50-per-Million-Token Multimodal Model for Coding and Agents

Unbiased has released Pareto, a multimodal composite model aimed at research, coding, and agentic workflows. The model offers a 262K context window and is priced at $2.50 per million input tokens and $7.50 per million output tokens via OpenRouter.

September 16, 2026
model release

Ex-OpenAI Researcher Launches Jev, an AI Model That Scores Options Instead of Generating Text

Startup TypeSafe AI has released Jev, a model built to score predefined answer options rather than generate text, claiming response times of 70 to 500 milliseconds. Co-founder Diogo Almeida, a former OpenAI researcher and InstructGPT co-author, says the model targets background classification tasks like sorting customer requests.

model release

Anonymous Provider Launches Union Alpha, a Free 262K-Context Multimodal Model on OpenRouter

A third-party provider using the alias 'Stealth' has released Union Alpha on OpenRouter, a multimodal model with a 262K context window, currently free to use during its preview period. The model's developer remains anonymous, and OpenRouter states it is not the model's owner or operator.

September 12, 2026
model release

Google Releases TimesFM-3, a 330M-Parameter Model That Forecasts Sales Using Weather and Discount Data

Google Research has released TimesFM-3, a 330-million-parameter time series forecasting model that predicts outcomes like sales by combining related variables, historical data, and known future events such as discounts or weather. The model claims top rankings on three benchmarks against Amazon's Chronos-2 and the Toto-2.0 family.

September 11, 2026
model releaseSakana Ai

Sakana AI Launches Fugu Ultra v2, a Multi-Agent Orchestrator With 1M-Token Context

Sakana AI has released Fugu Ultra v2, described as a learned multi-agent orchestration system rather than a single monolithic model. It offers a 1M-token context window, configurable reasoning effort, and pricing of $5 per 1M input tokens and $30 per 1M output tokens.

September 10, 2026
model releaseDeepSeek

DeepSeek V4.1-Flash Cuts KV Cache Memory by Up to 8x, Matches Opus 5 on Coding Benchmark

DeepSeek released V4.1-Flash, a 552-billion-parameter model built to slash the memory overhead of long-context AI agents. The model cuts GPU cache needs to roughly a quarter of its predecessor's and matches closed models from OpenAI and Anthropic on select coding benchmarks.

model releaseDeepSeek

DeepSeek Launches V4.1 Flash: Low-Cost MoE Model Claims to Beat V4 Pro

DeepSeek has released V4.1 Flash, a sparse mixture-of-experts model priced at $0.30 per 1M input tokens and $1.20 per 1M output tokens with a 1 million token context window. DeepSeek claims the model exceeds the larger V4 Pro on performance, speed, and task completion time.

September 9, 2026
model releaseSuno

Suno Releases v6 Music Model Family Trained on Licensed Data from Warner, BMG, Believe

Suno unveiled its v6 model family, trained on licensed data from Warner Music Group, BMG, and Believe, as the AI music startup continues fighting copyright lawsuits from Sony, Universal Music Group, and individual artists. The company plans to retire its older, non-licensed models entirely.

September 5, 2026
model releaseOpenAI

OpenAI Launches GPT-6 Astra With Half the Message Allowance of GPT-5.6 Sol

OpenAI has begun rolling out GPT-6 Astra to top-tier ChatGPT plans, the API, Azure, and AWS Bedrock. The model delivers roughly half the usage allowance of GPT-5.6 Sol across comparable plans, with Plus and Business users gaining access in the coming days.

September 3, 2026
model releaseOpenAI

OpenAI Releases GPT-6 Astra, First Model to Cross 'Critical' Cybersecurity Threshold

OpenAI has begun rolling out GPT-6 Astra, the first model to reach the company's internal 'Critical' cybersecurity threshold. Access is being phased, with companies in OpenAI's Daybreak cybersecurity program getting priority following added safeguards after a prior model containment breach.

model releaseOpenAI

OpenAI Launches GPT-6 Astra, Matches Claude Fable Pricing at $10/$50 per Million Tokens

OpenAI has begun rolling out GPT-6 Astra, priced at $10/million input and $50/million output tokens to match Claude Fable. The model claims a 99.9% score on ARC-AGI 3 using a custom harness and leads on security and long-context benchmarks, though it trails Fable on general intelligence rankings.

model releaseOpenAI

OpenAI Launches GPT-6 Astra, Says the Model May Already Qualify as AGI

OpenAI has released GPT-6 Astra, its most capable model yet, with benchmark scores the company says surpass GPT-5.6 Sol and Anthropic's Fable 5 models. President Greg Brockman called it a step into the 'AGI era,' though OpenAI acknowledges there's no agreed-upon threshold for that term.

model releaseOpenAI

OpenAI Releases Astra, Claims New Flagship Model Beats Rivals on Coding and Cybersecurity Benchmarks

OpenAI released Astra on Thursday, calling it its most capable and most aligned model yet. The model uses a reasoning technique called 'opaque recurrence' that critics say reduces visibility into its chain of thought.

model releaseInclusionai

InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters

InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.

changelog

Meta Releases Muse Spark 1.3, Cheapest Model in Its Performance Class at $0.55 Per Task

Meta has released Muse Spark 1.3, its fourth model in five months, with an xhigh tier available now and a more powerful max tier in limited preview. The model improves sharply on agentic benchmarks and costs $0.55 per index task—cheaper than any rival at the same performance level—but still trails Claude Fable 5.1 on most tests.

model release

Meta's Muse Spark 1.3 Claims #3 Global Ranking, Matches OpenAI's GPT-5.6-Sol on Coding Benchmarks

Meta Superintelligence Labs shipped Muse Spark 1.3, which the company claims ranks #3 globally on the Artificial Analysis Intelligence Index and matches OpenAI's GPT-5.6-Sol on coding and agentic benchmarks. The model is available now via Muse Code and Meta's API, with open weights and a follow-up model promised soon.

September 2, 2026
model release

Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window

Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.

model release

Google DeepMind Ships Gemini 3.8 Flash and a Cybersecurity Variant, Third Flash Release in Six Weeks

Google DeepMind released Gemini 3.8 Flash and a specialized cybersecurity variant, Gemini 3.8 Flash Cyber, its third Flash-tier launch in six weeks. Pricing stays at $0.75 per million input tokens and $3.75 per million output tokens, matching the prior 3.7 Flash release.

September 1, 2026
model releaseOpenAI

OpenAI's Astra Model Aces Cybersecurity Benchmark, Found Two Zero-Day Exploits Unassisted

OpenAI has disclosed new details on Astra, a forthcoming model the company says is the first to cross its 'critical cybersecurity threshold.' According to OpenAI, Astra scored a perfect result on ExploitBench and discovered two zero-day vulnerabilities in internal testing without human guidance.

model releaseOpenAI

OpenAI Says Upcoming Astra Model Is First to Cross 'Critical' Cybersecurity Risk Threshold

OpenAI says its upcoming Astra model is the first to cross its 'Critical' cybersecurity capability threshold, meaning it can discover and exploit unknown vulnerabilities without step-by-step human guidance. The company plans to release Astra soon but will restrict its advanced cyber capabilities to a vetted coalition of organizations.

changelogAnthropic

Anthropic Releases Fable and Mythos 5.1, Cuts Token Costs and Loosens Safeguard False Positives

Anthropic released Fable 5.1 and Mythos 5.1 on Tuesday, twinned models with reduced token costs and fewer false-positive safeguard triggers. Mythos remains restricted to cybersecurity and life sciences partners, while Fable is available now via cloud platforms and the Anthropic API.

model releaseAnthropic

Anthropic's Claude Fable 5.1 Launches on Amazon Bedrock and Claude Platform on AWS

Anthropic's Claude Fable 5.1 is now live on Amazon Bedrock and Claude Platform on AWS, improving on Fable 5 in reasoning, agentic coding, and long multi-step tasks. The model ships with new Enterprise Frontier Safeguards allowing zero data retention for eligible customers through December 2026.

model releaseAnthropic

Anthropic Releases Claude Fable 5.1, Cuts Agentic Workload Pricing Up to 45%

Anthropic has released Claude Fable 5.1, an upgrade to its top-tier Fable 5 model launched in June, alongside a restricted-access sibling called Mythos 5.1. The company claims the new model matches or beats Fable 5's performance while cutting costs by up to 45% on agentic workloads through reduced cache-read pricing.

model releaseInception

Inception Launches Mercury 2.5 Preview, a Diffusion LLM Claiming 1,107 Tokens/Sec

Inception released Mercury 2.5 Preview, a diffusion-based language model that generates tokens in parallel rather than sequentially, claiming throughput of 1,107 tokens per second on standard GPUs. The model is available on OpenRouter with a 260K context window and an 80% launch discount through September 8, 2026.

August 14, 2026
model release

Google DeepMind Ships Gemini 3.7 Flash, Closing Gap With Claude 4.8 and GPT-5.5

Google DeepMind has released Gemini 3.7 Flash, a new entry in its fast-tier model line that reportedly closes a performance gap that opened up under Gemini 3.5 and 3.6 Flash against Anthropic's Claude 4.8+ and OpenAI's GPT-5.5+ series. Full pricing and benchmark details have not yet been disclosed.

August 13, 2026
model releaseDeepSeek

DeepSeek Releases DeepSeek-V4-Pro-0813, a 1.7T-Parameter Model with DSpark Speculative Decoding

DeepSeek has released DeepSeek-V4-Pro-0813, a 1.7-trillion-parameter model that supersedes the DeepSeek-V4-Pro preview. The model adds a DSpark speculative decoding module and posts measurable gains on agentic and coding benchmarks, according to DeepSeek's technical report.