model release
50 articles tagged with model release
Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October
Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.
Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter
Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.
Google Releases Gemini 4 Argon to Cybersecurity Partners, Claims Wins Over GPT-6 Astra
Google has released Gemini 4 Argon, its next-generation flagship AI model, to a small group of cybersecurity partners as part of a phased rollout. The company claims the model outperforms OpenAI's GPT-6 Astra on several coding and knowledge-work benchmarks, though full specifications remain undisclosed.
Google DeepMind Releases Gemini 4 Argon, Expands Output Limit to 1M Tokens
Google DeepMind has released Gemini 4 Argon, a frontier model built for long-horizon reasoning with an industry-leading 1 million output token limit. The model is rolling out first to trusted cyber defenders through Google's Fairwind Program, with pricing set at $2 per million input tokens and $10 per million output tokens.
Google Announces Gemini 4 Argon, Its New Frontier Model With 1M Output Tokens
Google has announced Gemini 4 Argon as its new frontier model, featuring a 1M output token limit (up from 64K) and claimed leads on coding, cybersecurity, and automation benchmarks. The model is rolling out first to Google AI Ultra subscribers and paid API customers.
OpenAI Releases GPT-6.1 Sol: Mid-Tier Model with 1.1M Context at $2/$10 per Million Tokens
OpenAI has released GPT-6.1 Sol, an incremental upgrade to GPT-6 Sol positioned below flagship GPT-6 Astra in its GPT-6 lineup. The model features a 1.1M token context window, priced at $2 per 1M input tokens and $10 per 1M output tokens, with claimed improvements in factual accuracy and instruction-following on agentic tasks.
OpenAI Scraps Release of GPT-6.1 Astra Over Safety Concerns
OpenAI confirmed it will not release GPT-6.1 Astra after the model failed to meet internal safety and alignment standards. The decision follows renewed industry-wide calls, including from Anthropic, to slow the pace of frontier model development.
Anthropic's Claude Sonnet 5.5 Launches on Amazon Bedrock and Claude Platform on AWS
Anthropic's Claude Sonnet 5.5 is now available on Amazon Bedrock and Claude Platform on AWS, positioned as a faster, lower-cost model for well-scoped coding and document tasks. It pairs with the recently released Claude Opus 5.5, which handles higher-judgment work.
Anthropic Releases Claude Sonnet 5.5, Nearly Matching Opus 5.5 at Up to 30% Lower Cost Per Task
Anthropic has released Claude Sonnet 5.5, the second model in its Claude 5.5 family, delivering major coding and knowledge-work gains that approach flagship Opus 5.5 performance while using up to 30% fewer tokens per task. Per-token pricing stays unchanged at $2/$10 per million input/output tokens.
Anthropic Launches Claude Sonnet 5.5, Claims 30% Faster Performance at Lower Cost Than Predecessor
Anthropic has released Sonnet 5.5, the latest version of its mid-tier Claude model, claiming 30% faster performance and significantly lower token costs than its predecessor. The company says the model now outperforms Opus 5.5 on agentic coding tasks and carries cyber capabilities comparable to Opus 5.
Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context
Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.
AionLabs Launches Aion 3.5, a Multi-Model Storytelling System Built on GLM
AionLabs has released Aion 3.5, a collaborative multi-model system for roleplaying and storytelling built on the GLM model family. It offers a 262K token context window at $3 per 1M input tokens and $6 per 1M output tokens.
Upstage Releases Solar Mini 4: 35B MoE Model with 524K Context at $0.05/$0.20 per Million Tokens
Upstage has released Solar Mini 4, a compact mixture-of-experts model with 35B total parameters, 3B active parameters, and a 524K token context window. The model targets agentic workloads and is priced at $0.05 per 1M input tokens and $0.20 per 1M output tokens, a promotional 50% discount off standard rates.
OpenRouter Listings Surface for Three Unannounced OpenAI Models: GPT-6 Sol Pro, Luna, and Luna Pro
OpenRouter's model directory listed three new entries—GPT-6 Sol Pro, GPT-6 Luna, and GPT-6 Luna Pro—attributed to OpenAI, but no pricing, context window, benchmark data, or official confirmation from OpenAI has surfaced.
Anthropic Ships Claude Opus 5.5, OpenAI Launches GPT-6 Sol and Luna — All Cheaper Than Predecessors
Anthropic released Claude Opus 5.5 at $4/$20 per million input/output tokens, undercutting Opus 5's $5/$25 pricing while claiming better agentic coding scores. OpenAI countered with GPT-6 Sol ($2/$10) and GPT-6 Luna ($0.10/$0.50), both up to 50% cheaper than GPT-5.6's promotional rates.
OpenAI Launches GPT-6 Luna: Fast, Low-Cost Model With 1.1M Context Window
OpenAI has released GPT-6 Luna, the fast and cost-efficient entry in its new GPT-6 model family, featuring a 1.1M token context window and pricing starting at $0.10 per 1M input tokens. The model is positioned below GPT-6 Sol and GPT-6 Astra in OpenAI's tiered lineup.
OpenAI Launches GPT-6 Sol and GPT-6 Luna, Cutting API Prices 50% Versus GPT-5.6
OpenAI has released GPT-6 Sol and GPT-6 Luna, two new models that cost 50% less than their GPT-5.6 equivalents while claiming improved coding and computer-use performance. The models roll out today to ChatGPT Work and Codex for Plus, Pro, Business, Enterprise, and Edu users.
OpenAI's GPT-6 Sol and GPT-6 Luna Launch on Amazon Bedrock, Priced Below GPT-5.6
OpenAI has launched GPT-6 Sol and GPT-6 Luna on Amazon Bedrock, positioned below flagship GPT-6 Astra for recurring coding tasks and high-volume document processing respectively. Both models cost less per API call than their GPT-5.6 predecessors, though exact pricing figures were not disclosed.
Claude Opus 5.5 Launches on Amazon Bedrock, Anthropic's First Model in New 5.5 Family
Claude Opus 5.5, the first model in Anthropic's new Claude 5.5 family, is now live on Amazon Bedrock and Claude Platform on AWS. Anthropic claims the model does more with fewer tokens than Claude Opus 5, lowering average cost per task despite unchanged headline pricing tiers.
Anthropic Releases Claude Opus 5.5, Cuts Costs 40% While Matching Rival Fable 5.1
Anthropic has released Claude Opus 5.5, claiming performance parity with Claude Fable 5.1 at roughly 40% lower total operating cost than Opus 5. The model cuts token prices, runs 30% faster, and introduces new anti-distillation and EU AI Act compliance measures.
Xiaomi Lists Three New MiMo-V2.6 Models on OpenRouter: Pro, Flash, and Pro-UltraSpeed
Xiaomi has added three new entries to its MiMo model family on OpenRouter: MiMo-V2.6-Pro, MiMo-V2.6-Flash, and MiMo-V2.6-Pro-UltraSpeed. Full specifications, pricing, and benchmark results have not yet been disclosed.
Xiaomi Launches MiMo-V2.6-Pro-UltraSpeed: Same Quality, 10x Faster Output
Xiaomi's MiMo-V2.6-Pro-UltraSpeed is a fast-inference edition of the company's 1T-parameter flagship MiMo-V2.6-Pro, delivering roughly 10x the output speed at matching quality. It retains the 1M-token context window and native multimodal capabilities, priced at $4.35/$8.70 per 1M input/output tokens.
Xiaomi Launches MiMo-V2.6-Pro, a 1T+ Parameter Model with 1M-Token Context
Xiaomi has released MiMo-V2.6-Pro, a flagship foundation model exceeding 1 trillion parameters with a 1M-token context window and native multimodal support. The model is priced at $0.435 per 1M input tokens and $0.87 per 1M output tokens, targeting agentic and long-horizon tasks.
xAI Ships Grok 4.7, Cuts Price to $1.60/$4.80 per 1M Tokens With 500K Context
xAI has released Grok 4.7, the successor to Grok 4.6, listed on OpenRouter with a 500K token context window and pricing of $1.60 per 1M input tokens and $4.80 per 1M output tokens. The company claims improvements in long-running software engineering tasks, self-verification, and professional document drafting.
Unbiased Launches Pareto, a $2.50/$7.50-per-Million-Token Multimodal Model for Coding and Agents
Unbiased has released Pareto, a multimodal composite model aimed at research, coding, and agentic workflows. The model offers a 262K context window and is priced at $2.50 per million input tokens and $7.50 per million output tokens via OpenRouter.
Ex-OpenAI Researcher Launches Jev, an AI Model That Scores Options Instead of Generating Text
Startup TypeSafe AI has released Jev, a model built to score predefined answer options rather than generate text, claiming response times of 70 to 500 milliseconds. Co-founder Diogo Almeida, a former OpenAI researcher and InstructGPT co-author, says the model targets background classification tasks like sorting customer requests.
Anonymous Provider Launches Union Alpha, a Free 262K-Context Multimodal Model on OpenRouter
A third-party provider using the alias 'Stealth' has released Union Alpha on OpenRouter, a multimodal model with a 262K context window, currently free to use during its preview period. The model's developer remains anonymous, and OpenRouter states it is not the model's owner or operator.
Google Releases TimesFM-3, a 330M-Parameter Model That Forecasts Sales Using Weather and Discount Data
Google Research has released TimesFM-3, a 330-million-parameter time series forecasting model that predicts outcomes like sales by combining related variables, historical data, and known future events such as discounts or weather. The model claims top rankings on three benchmarks against Amazon's Chronos-2 and the Toto-2.0 family.
Sakana AI Launches Fugu Ultra v2, a Multi-Agent Orchestrator With 1M-Token Context
Sakana AI has released Fugu Ultra v2, described as a learned multi-agent orchestration system rather than a single monolithic model. It offers a 1M-token context window, configurable reasoning effort, and pricing of $5 per 1M input tokens and $30 per 1M output tokens.
DeepSeek V4.1-Flash Cuts KV Cache Memory by Up to 8x, Matches Opus 5 on Coding Benchmark
DeepSeek released V4.1-Flash, a 552-billion-parameter model built to slash the memory overhead of long-context AI agents. The model cuts GPU cache needs to roughly a quarter of its predecessor's and matches closed models from OpenAI and Anthropic on select coding benchmarks.
DeepSeek Launches V4.1 Flash: Low-Cost MoE Model Claims to Beat V4 Pro
DeepSeek has released V4.1 Flash, a sparse mixture-of-experts model priced at $0.30 per 1M input tokens and $1.20 per 1M output tokens with a 1 million token context window. DeepSeek claims the model exceeds the larger V4 Pro on performance, speed, and task completion time.
Suno Releases v6 Music Model Family Trained on Licensed Data from Warner, BMG, Believe
Suno unveiled its v6 model family, trained on licensed data from Warner Music Group, BMG, and Believe, as the AI music startup continues fighting copyright lawsuits from Sony, Universal Music Group, and individual artists. The company plans to retire its older, non-licensed models entirely.
OpenAI Launches GPT-6 Astra With Half the Message Allowance of GPT-5.6 Sol
OpenAI has begun rolling out GPT-6 Astra to top-tier ChatGPT plans, the API, Azure, and AWS Bedrock. The model delivers roughly half the usage allowance of GPT-5.6 Sol across comparable plans, with Plus and Business users gaining access in the coming days.
OpenAI Releases GPT-6 Astra, First Model to Cross 'Critical' Cybersecurity Threshold
OpenAI has begun rolling out GPT-6 Astra, the first model to reach the company's internal 'Critical' cybersecurity threshold. Access is being phased, with companies in OpenAI's Daybreak cybersecurity program getting priority following added safeguards after a prior model containment breach.
OpenAI Launches GPT-6 Astra, Matches Claude Fable Pricing at $10/$50 per Million Tokens
OpenAI has begun rolling out GPT-6 Astra, priced at $10/million input and $50/million output tokens to match Claude Fable. The model claims a 99.9% score on ARC-AGI 3 using a custom harness and leads on security and long-context benchmarks, though it trails Fable on general intelligence rankings.
OpenAI Launches GPT-6 Astra, Says the Model May Already Qualify as AGI
OpenAI has released GPT-6 Astra, its most capable model yet, with benchmark scores the company says surpass GPT-5.6 Sol and Anthropic's Fable 5 models. President Greg Brockman called it a step into the 'AGI era,' though OpenAI acknowledges there's no agreed-upon threshold for that term.
OpenAI Releases Astra, Claims New Flagship Model Beats Rivals on Coding and Cybersecurity Benchmarks
OpenAI released Astra on Thursday, calling it its most capable and most aligned model yet. The model uses a reasoning technique called 'opaque recurrence' that critics say reduces visibility into its chain of thought.
InclusionAI Releases Ling 3.0 Flash Fin, a Finance-Focused MoE Model with 5.1B Active Parameters
InclusionAI has released Ling 3.0 Flash Fin, a finance-specialized mixture-of-experts model built on Ling 3.0 Flash. The model activates 5.1B of its 124B total parameters and targets long-horizon investment planning tasks while retaining general reasoning, coding, and math capabilities.
Meta Releases Muse Spark 1.3, Cheapest Model in Its Performance Class at $0.55 Per Task
Meta has released Muse Spark 1.3, its fourth model in five months, with an xhigh tier available now and a more powerful max tier in limited preview. The model improves sharply on agentic benchmarks and costs $0.55 per index task—cheaper than any rival at the same performance level—but still trails Claude Fable 5.1 on most tests.
Meta's Muse Spark 1.3 Claims #3 Global Ranking, Matches OpenAI's GPT-5.6-Sol on Coding Benchmarks
Meta Superintelligence Labs shipped Muse Spark 1.3, which the company claims ranks #3 globally on the Artificial Analysis Intelligence Index and matches OpenAI's GPT-5.6-Sol on coding and agentic benchmarks. The model is available now via Muse Code and Meta's API, with open weights and a follow-up model promised soon.
Meta Releases Muse Spark 1.3 Contributor, a Low-Cost Multimodal Reasoning Model With 1M Context Window
Meta has released Muse Spark 1.3 Contributor, described as the cost-efficient contributor tier of its multimodal reasoning model line. The model offers a 1 million token context window at $0.10 per 1M input tokens and $0.20 per 1M output tokens, targeting experimentation and early-stage agentic workflows.
Google DeepMind Ships Gemini 3.8 Flash and a Cybersecurity Variant, Third Flash Release in Six Weeks
Google DeepMind released Gemini 3.8 Flash and a specialized cybersecurity variant, Gemini 3.8 Flash Cyber, its third Flash-tier launch in six weeks. Pricing stays at $0.75 per million input tokens and $3.75 per million output tokens, matching the prior 3.7 Flash release.
OpenAI's Astra Model Aces Cybersecurity Benchmark, Found Two Zero-Day Exploits Unassisted
OpenAI has disclosed new details on Astra, a forthcoming model the company says is the first to cross its 'critical cybersecurity threshold.' According to OpenAI, Astra scored a perfect result on ExploitBench and discovered two zero-day vulnerabilities in internal testing without human guidance.
OpenAI Says Upcoming Astra Model Is First to Cross 'Critical' Cybersecurity Risk Threshold
OpenAI says its upcoming Astra model is the first to cross its 'Critical' cybersecurity capability threshold, meaning it can discover and exploit unknown vulnerabilities without step-by-step human guidance. The company plans to release Astra soon but will restrict its advanced cyber capabilities to a vetted coalition of organizations.
Anthropic Releases Fable and Mythos 5.1, Cuts Token Costs and Loosens Safeguard False Positives
Anthropic released Fable 5.1 and Mythos 5.1 on Tuesday, twinned models with reduced token costs and fewer false-positive safeguard triggers. Mythos remains restricted to cybersecurity and life sciences partners, while Fable is available now via cloud platforms and the Anthropic API.
Anthropic's Claude Fable 5.1 Launches on Amazon Bedrock and Claude Platform on AWS
Anthropic's Claude Fable 5.1 is now live on Amazon Bedrock and Claude Platform on AWS, improving on Fable 5 in reasoning, agentic coding, and long multi-step tasks. The model ships with new Enterprise Frontier Safeguards allowing zero data retention for eligible customers through December 2026.
Anthropic Releases Claude Fable 5.1, Cuts Agentic Workload Pricing Up to 45%
Anthropic has released Claude Fable 5.1, an upgrade to its top-tier Fable 5 model launched in June, alongside a restricted-access sibling called Mythos 5.1. The company claims the new model matches or beats Fable 5's performance while cutting costs by up to 45% on agentic workloads through reduced cache-read pricing.
Inception Launches Mercury 2.5 Preview, a Diffusion LLM Claiming 1,107 Tokens/Sec
Inception released Mercury 2.5 Preview, a diffusion-based language model that generates tokens in parallel rather than sequentially, claiming throughput of 1,107 tokens per second on standard GPUs. The model is available on OpenRouter with a 260K context window and an 80% launch discount through September 8, 2026.
Google DeepMind Ships Gemini 3.7 Flash, Closing Gap With Claude 4.8 and GPT-5.5
Google DeepMind has released Gemini 3.7 Flash, a new entry in its fast-tier model line that reportedly closes a performance gap that opened up under Gemini 3.5 and 3.6 Flash against Anthropic's Claude 4.8+ and OpenAI's GPT-5.5+ series. Full pricing and benchmark details have not yet been disclosed.
DeepSeek Releases DeepSeek-V4-Pro-0813, a 1.7T-Parameter Model with DSpark Speculative Decoding
DeepSeek has released DeepSeek-V4-Pro-0813, a 1.7-trillion-parameter model that supersedes the DeepSeek-V4-Pro preview. The model adds a DSpark speculative decoding module and posts measurable gains on agentic and coding benchmarks, according to DeepSeek's technical report.