AI Model Changelog
Every version update across all major AI models.
October 2026
StepFun introduced Step 5 Preview, a 600B-parameter sparse MoE (27B active) with a 1.0M-token context window, aimed at agentic and professional knowledge work. It is priced at $1.00 input and $2.70 output per 1M tokens, with cache reads at $0.05.
Claude Haiku 5.5 replaces Haiku 4.5 as Anthropic's smallest model. Anthropic claims it costs around 75% less to run and is its fastest model to date.
TII released Falcon-ASR, a 1.6B-parameter speech recognition model for Arabic with a focus on the Emirati dialect, plus English, French, Spanish and Portuguese. TII claims a 20.92% average WER on the Open Universal Arabic ASR Leaderboard protocol and 22.73% WER on its internal Emirati evaluation.
Google released EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model on the Gemma 4 architecture under Apache 2.0. Google says quantized versions need about 191MB (text-only) to 567MB (full multimodal) of active RAM on a Pixel 11 Pro.
Musubi launched PolicyLM-1.7B, a 1.7B-parameter open-weights decision model for real-time content moderation. The company claims it applies plain-English policies in under 50 ms without retraining.
TII released Falcon-Emirati-7B, a 7B model built on Falcon-H1-Arabic and adapted to Emirati Arabic using native web data, MSA cultural material, and glossary-constrained synthetic data. TII claims it scores 84.83% on the Alyah benchmark, ahead of the Arabic and multilingual models it compared against.
Mistral released a public preview of Mistral Large 4, a 1T-parameter natively multimodal model with 49B active parameters, via the Mistral Studio API. Open weights are scheduled for release by the end of October 2026.
Nano Banana 2.1 is a Flash-tier Google image generation and editing model succeeding Nano Banana 2 and Nano Banana Pro. Google claims improvements in product recontextualization, mask- and ink-based editing, and factual accuracy, with 1K–4K output and a 66K context window.
Reka AI released a research preview of Rho-1, a 19B-parameter omni-model that handles text, images, video, and robot control actions in one network. Reka trained it on 320 H100 GPUs over about three months, using an inverse dynamics model to extract control signals from internet videos.
Reflection AI debuts Beam, its first frontier open-weight model: a text-only 501B-parameter MoE with 23B active parameters and a 1M token context window. The company claims reasoning performance on par with GLM-5.2 at 3-4x less inference compute, which is unverified.
Cloudflare released Clef, built on Qwen3.8-27B, and the smaller Clef-flash, built on Qwen3.5-9B. Both are decision models that output calibrated probabilities over predefined options, with text and image input.
Microsoft AI released MAI-Transcribe-2-Streaming, a 60-language real-time transcription model with first partial results in just over 100 ms. It launched alongside the MAI-Voice-2.1 and MAI-Voice-2.1-Flash text-to-speech models.
inclusionAI released Ling 3.1 Flash, a hybrid reasoning MoE model with 560B total and 25B active parameters and a 262K context window. It is listed as free on OpenRouter via NovitaAI, with no benchmark scores published.
Ai2 released AstaBrief 8B, an open-weights model fine-tuned from Qwen3-8B with SFT and DPO to generate cited scientific reports in one pass. It powers Asta's Fast mode, which Ai2 says is about 3.5x faster than the Claude-powered Thinking mode.
Unbiased released Pareto 26.10 Preview, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. It is a preview of the next Pareto version and may change without notice; pareto-26.9 is recommended for stable behavior.
Ideogram 4.5 is an image model that the company claims edits only user-specified regions while preserving the rest of the image, including after multiple edits. It outputs at native 2K across four per-image priced quality tiers.
September 2026
Gemini 4 Argon is Google's new frontier model, increasing output token capacity from 64K to 1M tokens and claiming leading scores on coding, cybersecurity, automation, and video understanding benchmarks. It replaces the canceled Gemini 3.5 Pro and is rolling out first to Google AI Ultra subscribers and paid API customers.
GPT-6.1 Sol upgrades GPT-6 Sol with fewer factual errors and improved reliability in respecting explicit restrictions during agentic tasks, while maintaining the same 1.1M context window and pricing.
Sonnet 5.5 replaces Sonnet 5, with Anthropic claiming 30% faster performance, lower token costs, and cyber capabilities now comparable to Opus 5, triggering the same cyber safeguards applied to Opus and Fable.
Mk1.5 succeeds Mk1 by adding audio input, expanding context from 33K to 37K tokens, and formalizing structured spatial/temporal annotations via a single annotation_format parameter.
Black Forest Labs released FLUX 3 Action, a 7B-parameter open robotics model built on FLUX 3 that predicts robot actions and environment changes from multi-camera video. The company claims a record success rate on the RoboLab-120 leaderboard at less than half the parameter count and up to 3.95x the speed of the previous best open model.
Fireworks Research released Ember-1, a reasoning model built on Kimi K3 that claims roughly 40% fewer reasoning tokens than its base model while maintaining comparable quality on internal evaluations.
Liquid AI released a 280M-parameter DSpark draft model for LFM2.5-VL-3B that enables speculative decoding, claiming decode speedups up to 3.13x on-device and 2.66x on H100 GPUs with no change to output quality. The drafter adds only 8.9% to the target model's parameter count and ships with day-one integrations for llama.cpp, MLX-VLM, and SGLang.
Alibaba released five new Qwen-Audio-3.1 models covering ASR, TTS, and real-time voice interaction, adding multi-speaker detection, emotion-aware generation, and single-pass audio synthesis, alongside price cuts of up to 95 percent for ASR, 85 percent for real-time, and 70 percent for TTS.
Aion Labs released Aion 3.5 Mini, a lower-cost version of Aion 3.5 built on the GLM model family, retaining the same 262K context window at roughly one-quarter the price.
Space Bunny Alpha is a new anonymous stealth model available free on OpenRouter, offering a 1M-token context window, multimodal input, and adjustable reasoning effort during its preview period.
Google released Gemini 3.8 Flash TTS and Flash-Lite TTS, adding generative voice design, voice replication with consent verification, and granular performance direction on top of the prior Gemini 3.1 Flash TTS model.
GLM-5.3-Prime is a high-speed inference variant of GLM-5.3, delivering claimed 1.5-2x output throughput while retaining the same 1M-token context window and full capabilities. It targets coding and agentic workloads with always-on reasoning at low, high, or max effort.
Upstage released Solar Mini 4, a 35B-parameter MoE model (3B active) with a 524K context window, replacing prior Solar Mini versions with lower cost and expanded context for agentic use cases.
NVIDIA released Nemotron 3 Diarization, a 100M-parameter open-weight model that expands speaker support from 4 to 8 speakers compared to its Streaming Sortformer predecessor, achieving a 14.72% DER and ranking #1 on Voice Arena's Diarization-Bench.
Xiaomi released MiMo-V2.6-Flash-RL, an efficiency-tier checkpoint of the MiMo-V2.6 series featuring native omnimodal input, 1M-token context, and a single mixed RL training run across coding, agent, visual, and cybersecurity domains. It joins the larger MiMo-V2.6-Pro-RL as part of the same model family release.
MiMo-V2.6-Pro-RL is a new 1.02T-parameter (42B active) omnimodal MoE model trained via a single mixed RL run, showing large benchmark gains over the prior MiMo-V2.5-Pro checkpoint, especially in cybersecurity and agentic tasks.
OpenAI released GPT-6 Luna, a fast and cost-efficient model in the GPT-6 family, offering a 1.1M token context window at $0.10/$0.50 per 1M tokens. It sits below GPT-6 Sol and GPT-6 Astra in OpenAI's tiered model lineup, with a Pro reasoning-mode variant also available.
Opus 5.5 lowers output pricing to $20 per million tokens from $25 and improves inference speed through reduced compute requirements, while also revising response style to prioritize key information and reduce jargon.
Xiaomi released MiMo-V2.6-Flash, an open-source 309B-parameter MoE model (15B active) expanding context from 256K to 1M tokens and adding native multimodal support compared to the prior MiMo-V2-Flash.
Xiaomi released a fast-inference edition of MiMo-V2.6-Pro built from the same 1T-parameter checkpoint, delivering roughly 10x the output speed at matching quality. Pricing is set at $4.35/$8.70 per 1M input/output tokens, exactly 10x the base model's rate.
Xiaomi released MiMo-V2.6-Pro, its flagship foundation model exceeding 1 trillion parameters, featuring a 1M-token context window and native multimodal capabilities. It is accompanied by MiMo-V2.6-Pro-UltraSpeed and MiMo-V2.6-Flash variants targeting speed and cost efficiency respectively.
Yandex released AliceAI-Foundation-80B-A3B-Base, a new 80B-parameter (3B active) hybrid MoE base model trained from scratch with a 262K-token context window. The model claims strong performance on Russian factual knowledge benchmarks and competitive math/reasoning scores against larger open-source MoE models.
Qwen-Image-2.1 is a new 7B-parameter open-weight image generation and editing model that adds native transparent image support, multi-reference image handling for up to 10 images, and faster inference via architecture changes and KV cache reuse.
Qwen released its first agent-focused multimodal model, Qwen3.8-Omni-Flash, claiming benchmark parity with Gemini 3.8 Flash on audio-video tasks at roughly 5-8x lower pricing. The release includes open-source Qwen-MM-Plugins and a real-time interaction tool called Qwen-Live Harness.
Initial release of Xing4.0-29B-A4B, a 29B-parameter (4B active) MoE model with 256K native context, marking the first model of this scale trained entirely on Huawei Ascend NPUs using MindSpore.
Initial preview release of Atria Dawn Preview, a 744B-parameter MoE agentic model built on the GLM-5.2 foundation model, with a 256K context window and text-only input.
PrismML released Ternary Bonsai 2 27B, a ternary-compressed reasoning model derived from Qwen3.8-27B that shrinks to ~8.5 GB while retaining a claimed 98.2% of base model performance. It supports a 262K-token context window, image understanding, tool calling, and thinks by default at 'xhigh' reasoning effort.
Z.ai released GLM-5.3-FlashX as a high-speed variant of GLM-5.3-Flash, using a 320B total / 18B active parameter hybrid attention architecture to deliver up to 200 tokens/second inference with a 1M-token context window.
Bonsai 2 27B compresses Alibaba's Qwen3.8 27B down to 5.9GB while retaining a claimed 98% of aggregate benchmark performance, up from 95% in the original Bonsai release in March. The model uses ternary weight quantization to achieve a 9x-10x reduction in memory footprint versus the uncompressed original.
Union Alpha is a new stealth multimodal model launched on OpenRouter with a 262K context window, free during its preview period. The developer remains anonymous; no independent benchmarks have been published.
Google released Gemini 3.8 Live and a reasoning-focused Extended Thinking variant for voice agents, pricing audio far below OpenAI's GPT-Live-1 while topping the Artificial Analysis Speech-to-Speech Leaderboard at 82.6 percent.
Google released Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, upgrading real-time voice dialogue with parallel reasoning, background tool execution, and visual grounding. Extended Thinking claims the #1 spot on Artificial Analysis' Speech to Speech Quality Index at 82.6.
ElevenLabs released Music v2.5, an updated version of its ElevenMusic model producing fuller, more natural-sounding tracks, especially in R&B, Hip-Hop, Rock, and orchestral genres according to blind testing. The model is now available via both the ElevenLabs app and API, with free (5 downloads/day) and Pro (400 downloads/month) tier options.
AllSpark released Iris-pro (397B parameters, built on Qwen3.5-397B-A17B) alongside the smaller Iris-mini (35B parameters, built on Qwen3.6-35B-A3B), both open-weight search agents trained via a reverse-engineered link-graph question pipeline and SFT-RL climbing. Iris-pro leads or ties among open-weight models in its size class on BrowseComp, BrowseComp-ZH, DeepSearchQA, and Humanity's Last Exam.
Inference.net released Schematron V2 Small, a 3B-parameter model specialized for extracting structured JSON from HTML pages using schema-defined instructions in response_format.
Inference.net released Schematron V2 Turbo, a 3B-parameter model specialized for high-volume HTML-to-JSON extraction with a 128K context window. Extraction rules must be defined via a JSON schema in response_format rather than through prompts.
Ling 3.0 Flash VL extends the text-only Ling 3.0 Flash MoE model with native image and video understanding, while InclusionAI claims further improvements to underlying language capabilities. The model retains the 124B total / 5.5B active parameter MoE architecture and 131K context window.
Sakana AI released Fugu Max, a cost-performance model that orchestrates a pool of open-weights and specialized models rather than functioning as a single monolithic network. It offers a 1M token context window with flat $2/$6 per-million-token pricing.
Fugu Ultra v2 is Sakana AI's higher-performance orchestration system, routing tasks across a fixed pool of open models with a 1M-token context window. It adds configurable reasoning effort levels, multimodal input, and built-in web search, priced at $5/$30 per 1M input/output tokens.
DeepSeek-V4.1-Flash introduces a Causal Encoder-Decoder architecture and Compressed Sparse Attention 2 that cut global KV cache to 890 bytes per token, roughly one-quarter of DeepSeek-V4-Flash. The 552B-parameter multimodal MoE model supports 1M-token context and activates only 8B/16B parameters per token.
PatchTST-FM-r2 upgrades the architecture from standard transformer layers to conformer blocks combining attention with temporal convolution, expands from 20 to 30 blocks, and adds overlapping patches with Hamming-window smoothing, improving zero-shot forecasting accuracy over predecessor PatchTST-FM-r1.
Suno replaced its prior model lineup with v6, a new family trained exclusively on licensed music data rather than the data underlying earlier models now facing copyright litigation. The release includes three variants — base, Wild, and Mini — with new editing, remixing, and multi-modal reference capabilities.
OpenAI released GPT-Image 2.5 as two API models, Sunburst and Flare, improving multi-turn instruction following, generation speed, and fidelity to subjects in reference photos. Sunburst is positioned for precision editing workflows; Flare for fast everyday generation.
OpenAI released ChatGPT Images 2.5, adding a Sketch feature that converts hand-drawn doodles into AI-generated images and an inline commenting tool for localized edits. OpenAI claims improved lighting, texture quality, multi-turn editing consistency, and up to 50% lower latency versus Images 2.0.
ChatGPT Images 2.5 improves detail, lighting, and texture quality while cutting generation latency by up to 50% versus Images 2.0, and adds a Sketch feature plus two new API models (Flare and Sunburst).
Qwen-Drive 1.0 extends Qwen3.5-4B with a perception module and a Planning Expert, enabling unified spatial understanding, route planning, and traffic dialogue in one model while reducing simulated off-road errors from 24% to 12%.
WeatherNext 3 replaces NWP-simulation training data with live geostationary satellite observations, enabling hourly forecast updates at up to 5km resolution — five times sharper than WeatherNext 2's 25km, 6-hour-increment forecasts.
GPT-6 Astra is OpenAI's newest flagship model, trained on over 100,000 GPUs and showing large gains over GPT-5.6 Sol in reasoning, coding, and cybersecurity benchmarks. It is priced 2.5x higher than Sol and is the first OpenAI model classified as 'critical' risk under the company's Preparedness Framework.
Google's Gemini 3.8 Flash debuted on OpenRouter with a 1-million-token context window and discounted pricing of $0.75/$3.75 per 1M input/output tokens, targeted at agentic and enterprise software engineering workloads.
World Labs released Atlas, its first omni-model trained from scratch on text, images, video, and 3D data to generate camera-controlled video, reconstruct 3D scenes from few images, and simulate environments for robotics. It is currently limited to an early-access partner program with no public pricing.
Meta released Muse Spark 1.3 Contributor, a low-cost tier of its multimodal reasoning model with a 1M token context window, priced at $0.10/$0.20 per 1M input/output tokens.
Meta released Muse Spark 1.3, a multimodal reasoning model with a 1M-token context window aimed at long-running agentic, multi-agent, and coding workflows. Audio input handling remains incompletely supported in this version.
Fable 5.1 upgrades Fable 5 with claimed improvements on coding and long-running problem-solving tasks, plus cache-read pricing cuts of up to 45% for agentic workloads. Anthropic also released Mythos 5.1, an identical model with restricted safeguards for cybersecurity and life sciences use.
Meta launched Muse Voice Transcribe, its first real-time audio perception model, combining speaker diarization, endpointing, and multilingual transcription in a single system. It is available via Meta's Mac app, Muse Code, and Model API at $3 per 1,000 audio minutes.
August 2026
Mercury 2.5 Preview succeeds Mercury 2 with a claimed 10+ point intelligence improvement and parallel token generation via diffusion, reaching claimed throughput of 1,107 tokens/sec.
IBM released Granite 4.2 8B, a dense 8B-parameter reasoning model with 131K context, three inference modes, and support for 12 languages, priced at $0.10/$0.15 per 1M tokens.
Fal post-trained Minimax's H3 video model for cost and quality, then optimized it for its proprietary inference engine, claiming up to 35x the speed of the official H3 endpoint. The result generates video faster than it can be watched, enabling live, chat-directed, infinite video streams.
Tencent introduced Hy4 preview, a mixture-of-experts model with 770B total and 49B active parameters, targeting coding agents and long-horizon tool-use tasks with a 1M token context window.
InclusionAI released Ling 3.0 Flash Fin, a finance-specialized MoE variant of Ling 3.0 Flash with 5.1B active parameters out of 124B total. The model adds a 262K-token context window and is tuned for investment workflows requiring long-horizon planning while retaining general reasoning, coding, and math abilities.
Gemini 3.5 Transcribe replaces Chirp 3 as Google's flagship speech-to-text model, claiming a 70% improvement in transcription latency and lower word error rates. It launches with function calling support, letting voice input trigger tasks like image generation across other Gemini models.
Alibaba released Qwen3.8 Flash, a multimodal reasoning model with a 1M-token context window, priced at $0.16 per 1M input and $0.47 per 1M output tokens via Alibaba Cloud International.
Z.ai released GLM-5.3-Flash, a native multimodal model with a 1M-token context window and a hybrid sparse-linear attention architecture aimed at coding and agent workloads. It launched with discounted pricing of $0.075/$0.25 per 1M input/output tokens through September 2026.
IBM released Granite 4.2 in 3B, 8B, and 30B sizes, trained on ~15 trillion tokens with context windows up to 512K tokens. The 8B and 30B variants received additional agentic RL training for tool use, code execution, and web search, all under Apache 2.0.
Qwen3.8-Flash-Next is an experimental preview of the architecture planned for Qwen4, introducing Qwen Sparse Attention, gated residual streams, n-gram embeddings, and a new optimizer recipe. It has 125B total parameters with 6B activated and supports up to 1 million tokens of context.
IBM released Granite 4.2, its first reasoning-focused LLM family in 3B, 8B, and 30B dense sizes, pre-trained on ~15T tokens with a five-phase pipeline and post-trained with multi-stage RL including agentic tool-use training for the 8B and 30B models.
IBM released two new 470M-parameter encoder-only Granite Speech models—one Apache 2.0, one non-commercial—replacing the prior LM-based architecture with a CTC-based design that is over 20x faster while achieving under 5% WER on public benchmarks.
Wan3.0 doubles maximum video length to 30 seconds compared to Wan2.5 and adds support for document and multimodal inputs including PDFs, PowerPoint files, images, video, and audio.
DeepSeek released an experimental vision-enabled variant of V4 Flash 0731, adding image understanding capabilities while matching the base model's text performance. The sparse MoE model uses 13B active parameters out of 284B total with a 1M token context window.
GEN-1.5 introduces in-context learning for robot manipulation, allowing a single short demo to be loaded as a 'physical prompt' instead of requiring dataset collection and training. Generalist reports 59% zero-shot success and 83% after brief fine-tuning across ten test tasks, though results are unverified.
Tencent released Hy-MT2-30B-A3B, a mixture-of-experts translation model with 30B total and 3B active parameters, supporting 33 language pairs and 5 Chinese dialect/minority-language pairs across five distinct translation workflows.
Ox Alpha is a new stealth reasoning model released anonymously through OpenRouter, offering a free preview with a 1 million token context window for coding and agentic use cases.
Qwen3.8-27B is a 27B dense vision-language model built on the Qwen3.5 architecture, offering native 262K context extensible to 1M tokens and configurable reasoning depth. This release provides FP8-quantized weights for efficient deployment across standard inference frameworks.
Qwen3.8-27B is a 27B-parameter dense vision-language model built on the Qwen3.5 architecture, improving on Qwen3.6-27B and Qwen3.7-Plus across coding, agentic, and multimodal benchmarks. It adds native image/video understanding, flexible thinking control, and a 262,144-token native context window extensible to 1 million tokens.
Writer released Palmyra X6, a post-trained variant of Z.ai's open source GLM-5.2, paired with an upgraded agentic harness, claiming up to 50% lower token costs for basic tasks.
MiniMax released Music 3, a new open-weight text-to-music model that generates complete, structurally coherent songs up to five minutes long from lyrics and detailed music descriptions.
Google released Gemini 3.7 Flash, a multimodal model with a 1,049K token context window aimed at fast agentic workflows and coding, priced at $0.38/M input and $1.88/M output tokens.
Alibaba released Qwen3.8-2.4T-A95B-FP8, an open-weight, FP8-quantized MoE checkpoint with 2.4T total and 95B activated parameters, serving as the foundation for the hosted Qwen3.8-Max API. The model extends context to 1M tokens and shows significant benchmark gains over Qwen3.7-Max on coding and agentic tasks, according to Alibaba.
Alibaba released Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter MoE model with 95B active parameters, built on the Qwen3.5 architecture with hybrid Gated DeltaNet and Gated Attention layers. It introduces reasoning_effort and preserve_thinking controls and claims Qwen-Max-class performance in an open-weight release for the first time.
xAI released Grok 4.6 with a 500K token context window and multimodal input support (text, image, file), now available via OpenRouter's API. The company claims it is their smartest model yet, with frontier performance on coding, knowledge work, and STEM tasks, though no independent benchmark scores were disclosed.
Initial public release of SL2T, Google DeepMind's massively multilingual sign-language-to-text model, deployed in Gboard and Live Transcribe on Pixel 11 for ASL-to-English translation.
LFM2.5-VL-3B upgrades Liquid AI's edge vision-language model with stronger screen/UI understanding, grounding, multi-image reasoning, and significantly improved function calling, while maintaining a 3.1B-parameter footprint for on-device deployment.
MAI Code 1.1 Flash improves on MAI-Code-1-Flash with a 25% token-efficiency gain and lower pricing, but trails DeepSeek-V4-Flash-0731 on Terminal Bench 2.1 and costs substantially more per output token.
Nemotron 3.5 Lightning succeeds Nemotron 3 Nano 30B A3B, gaining nine points on the Intelligence Index (15 to 24) and large jumps on agentic benchmarks (GDPval-AA v2 Elo up 334 points, Terminal-Bench v2.1 up from 7% to 24.3%) while becoming the fastest model in its comparison class at nearly 670 tokens per second.
ByteDance Seed released Seed-2.0-Code, a new coding-focused model with a 262K token context window and multimodal input support, designed for agentic coding and frontend development tasks.
Sakana AI released Namazu, a reasoning model fine-tuned from Kimi K2.6 for Japanese language and business contexts, with a 262K token context window.
NVIDIA released the full-precision BF16 reference weights for Nemotron 3.5 Lightning, a 30B-parameter hybrid Mamba-MoE-Attention model with 3B active parameters and up to 1M token context. This release is intended for fine-tuning and quantization rather than direct deployment, with a companion NVFP4 checkpoint available for optimized inference.
Upstage released Solar Pro 4, a text-to-text model with a 524K token context window aimed at agentic workflows, document processing, and coding, priced at $0.03/M input and $0.12/M output tokens.
NVIDIA released a new NVFP4-quantized checkpoint of Nemotron 3.5 Lightning, a 30B-total/3B-active hybrid Mamba-2/MoE/Attention model with 1M-token context, optimized for single-GPU deployment on DGX Spark or H100 hardware. The release includes post-training quantization and speculative decoding support to preserve BF16-level accuracy at lower inference cost.
Meta's first open-weight model since Llama 4 (spring 2025), a 30B-parameter agent model released under Apache 2.0, distilled from the larger Muse Spark model and optimized for local, on-device execution.
New 29.6B-parameter open-weight model with a built-in perception encoder, distilled from Muse Spark, designed for on-device agentic use with 4-bit quantization and speculative decoding.
GPT-5.6-Cyber is a new specialized variant of GPT-5.6 Sol trained for offensive security research, available exclusively through OpenAI's Daybreak Red access tier. It significantly outperforms its predecessor GPT-5.5-Cyber and standard safety-tuned models on sensitive security query completion and real-world exploit development.
Magpie TTS Multilingual adds Modern Standard Arabic, Korean, and Brazilian Portuguese, expanding coverage to 12 languages, and improves speech quality (lower CER, higher SSIM) on several existing languages including French and Spanish.
ByteDance Seed released Seed 2.1 Turbo, a multimodal model with a 262K token context window aimed at coding and agentic software delivery tasks. It is priced at $0.50 per 1M input tokens and $2.50 per 1M output tokens.
Meta released Muse Glimmer 30B, a dense open-weight model distilled from Muse Spark, optimized for agentic and coding workflows on consumer hardware with a 131K context window.
xAI launched Imagine Image 2.0 as a new Quality Mode with editing tools like Magic Wand, Multi-Ref Editing, and Smart Resize. The model ranks second globally on both the Image Edit and Text-to-Image Arena leaderboards, behind OpenAI's GPT-Image-2.
FLUX 3 Video moves from limited access to general availability via the BFL API and select partners, adding native audio, multi-scene generation, and 14+ language lip-sync. BFL claims the model outperforms Seedance 2.0, Gemini Omni Flash, and Minimax H3 on its internal Elo benchmarks.
Meta released Muse Spark 1.2 alongside its first coding agent, Muse Code, with the two trained together for improved coding performance over the standalone Muse Spark 1.1.
K-EXAONE 2.0 scales up from its predecessor's 236B/23B parameter MoE architecture to 750B/37B parameters via upcycling, adding long-context retrieval strength, expanded multilingual coverage (6 to 10 languages), and improved safety benchmark performance.
Mistral released Shieldstral, a 3B open-weights multimodal safety classifier under Apache 2.0 that reframes moderation as policy-adaptive question answering, claiming performance matching guard models up to 7x its size.
Liquid AI released LFM2.5-2.6B, a 2.6B-parameter open-weight model trained via supervised fine-tuning, teacher specialization, on-policy distillation, and agentic RL for local tool-use agents. The model claims competitive performance with models up to 4x its size on instruction-following, tool-use, and agentic benchmarks while running at 220 tok/s on Apple M5 Max hardware.
Qwen3.8-Max scales to 2.4 trillion total parameters (95 billion active) and is built on the Qwen3.5 architecture, with training focused on sustaining autonomous, multi-day agentic tasks rather than single-turn responses. It is the first Qwen-Max model to be open-weighted, with weights arriving on Hugging Face and ModelScope next week.
Initial v1.0 release of NemotronLabs VoiceChat 11B, NVIDIA's first open full-duplex speech model combining a Fast Conformer encoder, Nemotron Nano V2 9B LLM backbone, and TTS decoder into one architecture. The model adds live tool-calling support during natural conversation, a first among open full-duplex voice models according to NVIDIA.
MiniMax released open weights for H3, a 33B-parameter video model that ranks first in Video Editing and second in Text-to-Video on Artificial Analysis. The 2K resolution module and H3-Context-IR remain closed, capping local generation at 768p.
DeepSeek introduced a rolling 'latest' alias for its V4 Flash model family, currently offering a 1,049K token context window at $0.09/M input and $0.18/M output tokens.
Seedance 2.5 extends maximum clip length to 30 seconds (up from shorter limits in 2.0) and expands reference input support to 30 images, 10 videos, and 10 audio files for multi-character scene construction. ByteDance also claims improved visual fidelity in textures, lighting, and skin detail.
July 2026
DeepSeek released V4 Flash 0731, a re-post-trained sparse MoE checkpoint with 13B active/284B total parameters and a 1049K token context window, priced at $0.14/M input and $0.28/M output.
OpenAI cut GPT-5.6 Terra pricing by 20% to $2/$12 per million input/output tokens and GPT-5.6 Luna pricing by 80% to $0.20/$1.20 per million tokens, three weeks after launch. GPT-5.6 Sol pricing is unchanged.
Gemini Robotics ER 2 upgrades from ER 1.6 with continuous video-based progress tracking, faster moment-finding, multi-robot collaboration, and improved safety behavior around humans. It is now available via the Gemini API and Google AI Studio, with private preview on the Gemini Enterprise Agent Platform.
Gemini Robotics 2 extends the original Gemini Robotics system beyond arm-and-hand manipulation to full-body humanoid control, using a multi-model architecture and a new safety benchmark called ASIMOV-Agentic.
Thinking Machines Lab released Inkling Small, an open-weight multimodal MoE model with 12B active parameters out of 276B total and a 524K context window.
GPT Transcribe improves word error rate to 3.31 percent from GPT-4o Transcribe's prior score, and OpenAI cut pricing 25 percent to $0.0045 per minute. The model still trails ElevenLabs Scribe v2 (2.3%), Google Gemini 3 Pro (2.9%), and Mistral Voxtral Small (3%) on the AA-WER benchmark.
Lyria 3.5 improves melodic complexity, lyric quality and prompt adherence, vocal realism and pronunciation, and adds finer control over tempo and track duration, according to Google.
Ai2 launched the OlmoEarth Platform, infrastructure for fine-tuning, evaluating, and running large-scale inference with its OlmoEarth Earth observation foundation models. The platform claims continent-scale processing in about a day using a three-stage CPU/GPU/CPU pipeline and a custom satellite imagery metadata index.
Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoder models converted from LFM2.5 decoder backbones, offering 8,192-token context with claimed 3.7x faster CPU inference than ModernBERT-base at long context.
Alibaba released Qwen3.7 Flash, a vision-language reasoning model with a 1M token context window, priced at $0.03 per 1M input tokens and $0.13 per 1M output tokens.
Microsoft's first cybersecurity-specialized model, designed to find vulnerabilities in complex codebases and power the MDASH harness. Launched alongside Perception, an agentic platform for automated security operations.
Composer 2.5 is Cursor's low-cost worker model, built on Kimi K2.5 according to Cursor founder Michael Truell, and priced at $0.50/$2.50 per million input/output tokens. Cursor claims it performs at a level comparable to Opus 4.7 and GPT-5.5 in agentic coding tasks when paired with a frontier planner model.
Claude Opus 5 succeeds Opus 4.8 as Anthropic's flagship model, with the company claiming performance comparable to rival Fable 5 at roughly half the cost. Anthropic says it leads on coding and business-workflow benchmarks while requiring about 85% less intervention than Fable 5.
Google introduced Gemini Omni Flash Preview, a native multimodal model that generates 720p videos with synchronized audio from text, image, and video inputs. This is a new preview release, not an update to an existing model.
Fugu Ultra v1.1 claims performance gains of up to 7.9 points over v1.0, driven largely by improvements on ProgramBench and TerminalBench 2.1, and adds a Claude Code-compatible terminal endpoint. Sakana claims the update outperforms Anthropic's Fable 5 despite Fable 5 not being part of the router's model pool, though all figures are unverified.
Google expanded access to Gemini Spark, an agentic assistant powered by Gemini 3.5, to all US Google AI Pro subscribers and most Google AI Ultra subscribers globally. Free-tier users and Ultra subscribers in the EEA, Switzerland, UK, and Nigeria are excluded from this rollout.
InclusionAI released Ling-3.0-flash, a 124B-parameter MoE model activating ~5.1B parameters per token, offered free via OpenRouter with a 262K context window.
Flux 3 adds native audio generation to videos up to 20 seconds long and introduces a unified Self-Flow architecture trained jointly on images, video, audio, and actions. A companion robotics model, Flux-mimic, is being tested at Audi.
Gemini Nano 4, built on Gemma 4, launches on Samsung's Galaxy Z Fold 8, Fold 8 Ultra, and Flip 8, bringing improved multimodal understanding and expanded language support as the on-device foundation for Google's new Gemini Intelligence feature tier.
Laguna S 2.1 launches: Poolside's 118B open-weight MoE (8B active) for agentic coding, 1M context, OpenMDW-1.1 license.
Qwen-Image-3.0 expands prompt length to 4,500 tokens and improves text rendering fidelity down to ten pixels, targeting practical layouts like infographics and newspaper pages rather than purely aesthetic images. It is currently invite-only via API and unlikely to be released with open weights, unlike its predecessor.
Initial release of 1.14B parameter multilingual embedding model derived from Ministral-3B through two rounds of structured pruning and distillation.
First multimodal Qwen model above 1 trillion parameters with 2.4T parameters total. Claims improved coding and office work capabilities over Qwen3.7-Max.
FLUX.2 represents a major architecture update with a new 32B parameter model combining Mistral-3 24B VLM with rectified flow transformer, new FLUX.2-VAE trained from scratch, 4MP editing support, and 10-image multi-reference capability.
Initial release of LongCat 2.0, a sparse MoE model with 1.6T total parameters and over 1M token context window, optimized for coding and agentic tasks.
Kimi K3 represents a major release from Moonshot AI, claiming frontier-level performance while remaining open source. Independent analyses from Arena.ai and Vals AI suggest competitive performance with flagship frontier models.
Kimi K3 achieves top ranking in Arena's front-end coding capability benchmark, positioning as an open-source alternative to closed U.S. models at 50% lower pricing than GPT-5.6 Sol.
Cosmos 3 Edge is Nvidia's edge-optimized world model for physical AI applications, following the Cosmos 3 release in May 2026. Designed for robots and vision AI agents requiring real-time environmental perception.
New flagship 8B embedding model ranks #1 on RTEB with 78.5% score. Includes efficient 1B variants with 27% error reduction and Blackwell-optimized NVFP4 deployment option.
Initial release of Inkling, a 975B-parameter MoE model with 41B active parameters trained on 45 trillion tokens. Apache-2.0 licensed for commercial use and fine-tuning.
First public model release from Thinking Machines Lab. Open-weight mixture-of-experts architecture designed for enterprise customization rather than frontier performance.
Version 2.5 update to the KAT-Coder-Air lightweight coding model line with 256K context window support.
Gemma 4 E2B is a new variant optimized for Pixel 10's Tensor G5 TPU, enabling fully offline multimodal AI capabilities including chat, image recognition, and audio transcription.
Initial release of Bonsai 27B, claiming to be the first 27-billion parameter model capable of running on-device on iPhone with approximately 4GB memory footprint using 1-bit and ternary quantization.
GPT-5.6 Sol introduces adjustable reasoning effort levels across three model sizes, with approximately five to six settings per size. The Ultra setting uses four subagents to accelerate work at Max-level effort.
Muse Spark 1.1 is an updated version of Meta's foundation model, specifically designed for agentic and coding workloads. Released three months after the original Muse Spark.
Version 2.5 release of KAT-Coder-Pro coding model with 256K context window and updated pricing structure through OpenRouter.
GPT-5.6 Luna introduces a cost-optimized model in the GPT-5.6 series with 1M context window at $1/$6 per 1M tokens. Released July 9, 2026 with February 2026 knowledge cutoff.
GPT-5.6 Sol Pro launches as a reasoning-optimized variant of GPT-5.6 Sol, automatically applying extended reasoning mode for higher-quality responses on complex tasks.
New reasoning-optimized variant of GPT-5.6 Terra with pro mode enabled by default for higher-quality responses on complex tasks.
GPT-Live-1 introduces full-duplex voice architecture allowing ChatGPT to speak and listen simultaneously, with ability to conduct web searches during conversations. Available in full version for paid users and mini version for free users.
Compressed variant of Nemotron-3-Super reduced from 120.7B/12.8B active to 75.3B/9.3B active parameters using Iterative Puzzle framework. Achieves 2× throughput and 8× H100 concurrency while maintaining benchmark accuracy.
Initial release of unified audio-text model built on Nemotron-Cascade-2-30B-A3B backbone with extended vocabulary for discrete audio tokens and audio encoder.
GPT-5.6 Luna Pro is a reasoning-enhanced variant of GPT-5.6 Luna with reasoning.mode set to pro for higher-quality responses on complex tasks. Features 1M context window and February 2026 knowledge cutoff.
GPT-5.6 Terra introduces a balanced mid-tier option in the GPT-5.6 series with 1M context window at $2.50/$15 per million tokens, positioned between flagship Sol and cost-efficient Luna tiers. Targets everyday coding, reasoning, and agentic tasks with February 2026 knowledge cutoff.
GPT-Live-1 replaces OpenAI's turn-based voice model with full-duplex architecture that processes audio input and output simultaneously. The model integrates with GPT-5.5 for reasoning tasks and adds real-time translation.
DeepSeek-V4-Flash introduces a 284B parameter MoE architecture with 13B activation, 1M context window support, and three distinct reasoning modes. Unsloth provides optimized GGUF quantizations with Q8 at 162GB.
Laguna XS 2.1 improves on XS.2 with a 5.4% gain on SWE-bench Multilingual (63.1% vs 57.7%) and enhanced terminal-style task performance, while maintaining the 33B total parameter count with 3B activated per token.
Initial release of Robostral Navigate, an 8B embodied navigation model trained entirely in simulation using 400,000 trajectories across 6,000 scenes. Achieves 76.6% R2R-CE validation unseen success using only single RGB camera.
First release of Muse Image, Meta's first AI image generation model from Meta Superintelligence Labs. Introduces Instagram account prompting, claimed QR code generation, and integration across Meta AI app, Instagram, and WhatsApp.
Initial release of 2B parameter Conformer-based ASR model achieving 25.87% average WER on Open Universal Arabic ASR Leaderboard, outperforming models up to 30B parameters.
Initial release of Aion-3.0-Mini, a multi-model collaborative system built on DeepSeek for roleplaying and storytelling applications.
M2.5 is trained specifically for agent-native execution using reinforcement learning on agentic scaffolds, with emphasis on tool-calling, multi-step task decomposition, and long-horizon coding tasks.
Initial release of Nex-N2-Mini, the smaller model in the Nex-N2 series with 262K context window and open-source weights.
First release of Nemotron-Labs-TwoTower, a block-wise diffusion model built on the Nemotron-3-Nano-30B backbone. Uses dual-tower architecture to generate blocks of tokens in parallel, claiming 2.42× speedup while retaining 98.7% of baseline quality.
Leanstral 1.5 delivers major performance upgrades over the original Leanstral, saturating miniF2F at 100%, solving 587/672 PutnamBench problems, and achieving state-of-the-art results on FATE-H (87%) and FATE-X (34%) while reducing cost to $4 per problem versus $300+ for competitors.
Redeployed version with enhanced safety classifiers that automatically block cybersecurity tasks and revert to Opus 4.8. Available with restricted usage limits after month-long government-mandated takedown.
Initial release of Portugal's first national AI model, built on EuroLLM-9B foundation with European Portuguese datasets, multimodal capabilities, and expanded context window.
June 2026
Google released Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite), positioned as their fastest and cheapest image generation model optimized for velocity and scale.
Google DeepMind released Nano Banana 2 Lite as the fastest, most cost-efficient image model in the Nano Banana family, designed for high-volume developer workflows. It delivers 4-second text-to-image generation at $0.034 per 1K-resolution image, replacing the original Nano Banana model.
Introduces 1M token context, 128K max output, removes sampling parameters (temperature, top_p, top_k), and implements new tokenizer generating ~30% more tokens than Sonnet 4.6 for English text.
Nano Banana 2 Lite replaces the original Nano Banana as Google's fastest and cheapest image generator, generating images in 4 seconds at under $0.04 per thousand images. The model prioritizes speed and cost over quality for high-volume developer pipelines.
New lightweight image generation model optimized for speed and cost, generating 1K images in approximately 4 seconds while maintaining character consistency and editing capabilities of the Nano Banana family.
Claude Sonnet 5 is Anthropic's first Sonnet model of their latest generation, delivering near-Opus intelligence at Sonnet pricing. Key improvements include stronger multi-step reasoning, improved code navigation, and enhanced agentic reliability for production workflows.
Initial release of Ornith-1.0, an MIT-licensed agentic coding model built on Gemma 4 and Qwen 3.5, available in 9B, 31B, 35B MoE, and 397B MoE variants.
Major release introducing hybrid attention architecture with 90% KV cache reduction, 1M token context window, three reasoning modes, and trained on 32T+ tokens. Includes both Pro (1.6T params) and Flash (284B params) variants.
LFM2.5-230M is a new compact model in the LFM family, built on LFM2 architecture with extended pre-training and reinforcement learning. It features 230M parameters, tool use capabilities through distillation from LFM2.5-350M, and optimized edge inference performance.
DeepSeek-V4-Fable is a distilled variant of Claude-5-Fable built on DeepSeek-V4-Flash, trained specifically for offensive security research using 80,000 CTF trajectories and GRPO reinforcement learning.
First release of Un-0, an image generation model running on simulated oscillator hardware. Produces results comparable to diffusion models like Stable Diffusion but runs on software simulation of non-existent hardware.
Third behavioral update to GPT-5.5 Instant focused on conversational quality, with improved intent recognition, constraint handling, and recommendation capabilities.
Initial release of Qwen-AgentWorld-35B-A3B, a language world model trained through CPT, SFT, and RL stages specifically for agentic environment simulation across seven unified domains.
Initial release of Fugu Ultra, the higher-performance model in Sakana AI's Fugu family using learned multi-agent orchestration rather than a monolithic architecture.
OCR 4 adds bounding boxes, block classification, and inline confidence scores to document extraction. The model expands language coverage to 170 languages and introduces single-container self-hosted deployment.
Updated version achieves 85.6% on CyberGym benchmark, surpassing Anthropic's Mythos 5. Released alongside expanded international partnerships.
Initial release of Krea 2 family with Raw (base) and Turbo (post-trained) checkpoints. Turbo variant includes fine-tuning and distillation for faster 8-step generation.
Initial release of Unlimited-OCR, a 3B parameter OCR model building on Deepseek-OCR with support for single-page and multi-page document parsing.
Initial release of Laguna M.1, a 225B parameter MoE model with 23B active parameters designed for agentic coding tasks, featuring 256 experts and 262K token context window.
First major release of Mistral Large 3, a sparse MoE model with 675B total parameters. Includes multimodal capabilities, multilingual support, and ranks #2 on LMArena among OSS non-reasoning models.
Initial release of Voxtral TTS, Mistral's first text-to-speech model with 4B parameters supporting 9 languages and voice cloning from minimal audio samples.
First release of Leanstral, a specialized model for Lean 4 proof assistant with 6B active parameters from 120B total. Apache 2.0 licensed with free API endpoint and Mistral Vibe integration.
Devstral 2 powers Mistral Vibe 2.0 with support for custom subagents, multi-choice clarifications, and slash-command skills. API access moved from free to paid pricing.
Mistral OCR 3 achieves 74% win rate over OCR 2 with major improvements in handwriting, forms, scanned documents, and complex tables. Priced at $2 per 1,000 pages ($1 with Batch API).
First release of Mistral's code-specialized embedding model with 8192 token context, flexible dimensions, and optimized performance on code retrieval benchmarks including SWE-Bench Lite.
Mistral Medium 3 is a new model introduced alongside Le Chat Enterprise, positioned between Mistral Small and Large tiers. No technical specifications or benchmarks were disclosed.
Initial release of Mistral OCR API with multimodal document understanding at $1 per 1,000 pages. Claims state-of-the-art accuracy on internal benchmarks across math, tables, and multilingual content.
Pixtral Large is Mistral's new multimodal model, significantly more powerful than Pixtral 12B for document and image understanding. Powers le Chat's document analysis capabilities.
Gemini 3.1 Flash Image (Nano Banana 2) introduces image generation and editing capabilities with claimed Pro-level quality at Flash model speed and cost efficiency.
Initial release of Nano Banana Pro extends the original Nano Banana model with significantly improved multimodal reasoning, real-world grounding, and professional-grade image generation capabilities including 2K/4K outputs.
NVIDIA quantized Google DeepMind's DiffusionGemma 26B A4B IT from 16-bit to 4-bit (NVFP4) using Model Optimizer, reducing memory requirements while maintaining benchmark performance within 1% of the full-precision baseline.
Initial release of Cohere's North family. First agentic coding model from Cohere featuring sparse MoE architecture with 30B total parameters and 3B active.
Gemma 4 family launches on Amazon Bedrock with three variants: 31B dense model, 26B-A4B mixture-of-experts model, and E2B compact model. All variants support multimodal input and built-in reasoning mode.
Initial release of FastContext-1.0 repository explorer family with 4B and 30B parameter variants trained via supervised fine-tuning and reinforcement learning.
DiffusionGemma converts the existing Gemma-4-26B-A4B model into a text diffusion model using under 10% of the original training token budget, rather than training from scratch. It trades some quality on standard benchmarks for significantly higher generation speed and new self-correction capabilities.
Kimi K2.7 Code builds on K2.6 with substantial improvements in real-world long-horizon coding tasks and 30% reduction in thinking token usage. Added experimental video input support and improved performance across coding and agentic benchmarks.
Initial release of M3 with 428B parameters, native multimodal training, and MiniMax Sparse Attention enabling 1M context with 15× decode speedup over M2.
Third-generation Apple Foundation Model with 20B parameters using sparse activation. First on-device Apple model to exceed 3B parameters and support native multimodal capabilities.
Google releases DiffusionGemma 26B as open-weight model under Apache 2 license, bringing diffusion-based text generation to production with 500+ tokens/second inference speed.
Initial release of DiffusionGemma, a discrete diffusion-based text generation model built on Gemma 4 26B A4B MoE architecture with encoder-decoder design for parallel token generation.
Initial release of North Mini Code 1.0, a 30B-parameter sparse MoE model with 3B active parameters, trained specifically for agentic coding tasks with tool-use capabilities.
Fable 5 was shut down by government order on June 12, 2026, just 3 days after launch. Expected to return with potential access restrictions after government security review.
New speech-to-speech translation model with continuous audio generation across 70+ languages, expanding from Google Meet's previous 5-language support to enable 2000+ language combinations without requiring English as intermediary.
Third-generation AFM introduces flash-based inference architecture enabling 20B-parameter on-device model. Uses expert pruning to load subset of parameters into memory while keeping full model in flash storage.
First public release of Anthropic's Mythos model line. Introduces extended-execution code generation with up to 12-hour continuous operation on complex specifications.
Apple's most advanced cloud-based AI model, running on Nvidia GPUs in Google cloud infrastructure with privacy guarantees. Refined using Google Gemini frontier model outputs.
Initial release of Nex-N2-Pro, a 397B parameter MoE model with 17B active parameters and 262K context window, available free via OpenRouter.
Initial release of Nemotron 3 Ultra, a 550B-parameter MoE model with 55B active parameters, hybrid Transformer-Mamba architecture, and 1M context window designed for agentic workflows.
First open-weight release from Ideogram featuring 9.3B parameters, trained from scratch with structured JSON prompting, native 2K support, and state-of-the-art text rendering capabilities.
First release of Nemotron-3-Ultra, NVIDIA's 550B parameter (55B active) frontier model with hybrid LatentMoE architecture combining Mamba-2, MoE, and Attention layers. Features 1M token context and toggle-able reasoning capabilities.
Nemotron 3.5 adds custom policy enforcement, unified multimodal evaluation, auditable reasoning traces, and releases the training dataset. Built on Gemma 3 4B IT with 128K context window.
Nemotron 3.5 ASR expands the English-only Nemotron 3 ASR to support 40 language-locales from a single 600M-parameter checkpoint with cache-aware streaming architecture.
Initial release of Nemotron-3-Ultra, a 550B parameter model trained December 2025-April 2026 with hybrid LatentMoE architecture, 1M token context, and configurable reasoning capabilities.
First mid-sized Gemma model with native audio support, eliminates multimodal encoders for direct vision and audio processing through LLM backbone. Requires only 16GB RAM for local inference.
Qwen3.7 Plus introduces multimodal capabilities to the Qwen3.7 series with text and image input support, featuring a 1 million token context window at competitive pricing.
Gemma 4 introduces encoder-free architecture in the 12B Unified model, processing all modalities directly through a single decoder-only transformer. The family spans five models from 2.3B to 30.7B parameters with extended context windows up to 256K tokens.
First open-weight release of Ideogram's text-to-image model with 9.3B parameters, structured JSON prompting, and native 2K resolution support.
Gemma 4 12B is a new 12-billion-parameter multimodal model designed for local inference on consumer laptops. Google claims it matches the performance of their 26-billion-parameter mixture-of-experts model while running on devices with 16GB RAM.
Initial release of MAI-Thinking-1, a 35B parameter reasoning model trained on commercially licensed data without third-party model distillation. Available to select early partners.
Initial release limited to Project Glasswing partners. Public Mythos-class models planned for release within weeks of June 2, 2026.
Holo3.1 adds mobile automation support, native function-calling, and quantized checkpoints for local inference. AndroidWorld performance improves from 67% to 79.3%, with 2× end-to-end speedup on DGX Spark.
Initial release of Mellum2-12B-A2.5B-Thinking, a reasoning-augmented model trained with supervised fine-tuning and reinforcement learning with verifiable rewards on math-heavy data.
Initial release of Cosmos3-Nano, a 16B-parameter omnimodal world model for Physical AI applications with 256K context window and support for generating video, audio, images, and robot actions from multimodal inputs.
Anthropic expands restricted preview to 150 additional organizations across 15+ countries, focusing on critical infrastructure sectors including utilities and healthcare. Public release remains delayed pending safety measures.
Initial release of Microsoft's first advanced reasoning model, trained from scratch on proprietary data without distillation from other models.
Initial GGUF release of Step-3.7-Flash with seven quantization variants from BF16 (394GB) to IQ3_XXS (76GB), all optimized for local deployment on consumer hardware with 128GB memory.
First release of Cosmos 3, a unified omni-model that combines video generation, physical reasoning, and action prediction in a single Mixture-of-Transformers architecture, replacing the previous fragmented Cosmos model suite.
M3 introduces MiniMax Sparse Attention to enable 1M-token context at approximately 1/20th the compute cost of previous generation. Native multimodal training on interleaved data with interactive user-simulator tuning.
May 2026
Initial release of Cosmos 3, NVIDIA's omnimodal world foundation model platform for Physical AI, featuring 64B-parameter variants with Mixture-of-Transformers architecture supporting video, image, audio, and robot action generation.
Initial release of Cosmos3-Super-Text2Image as part of NVIDIA's Cosmos3 omnimodal world model collection, featuring 64B parameters and Mixture-of-Transformers architecture for Physical AI applications.
Initial release of Step-3.7-Flash, a 198B-parameter sparse MoE vision-language model with 256K context window, three reasoning levels, and production-focused architecture delivering 400 tokens/sec.
Opus 4.8 shows substantially lower rates of misaligned behavior compared to 4.7, with approximately 4x improvement in catching code flaws. Fast mode API pricing reduced by 3x.
Fast-mode variant of Claude Opus 4.8 offering higher output speed at 2x pricing compared to the standard version.
Initial release of Mistral Saba, a 24B-parameter model specialized for Arabic and South Asian languages including Tamil, designed for regional deployment.
First reasoning model release from Mistral AI with 24B parameters and native multilingual chain-of-thought capabilities. Open-sourced under Apache 2.0 license.
New release of Devstral Medium achieving 61.6% on SWE-Bench Verified. Companion release of Devstral Small 1.1 (24B parameters) scores 53.6% and is open-sourced under Apache 2.0.
Initial release of Voxtral family with 24B (Small) and 3B (Mini) variants under Apache 2.0. Features 32K context, native multilingual support, and direct function-calling from speech.
Mistral Large 3 is Mistral's first MoE model since Mixtral, featuring 675B total parameters with 41B active. Released with both base and instruction-tuned versions under Apache 2.0, with multimodal image understanding and optimized inference support.
Codestral 25.08 improves code completion with 30% higher acceptance rates, 10% better retention, and 50% fewer runaway generations. Chat mode gains 5% on instruction-following and code benchmarks.
Mistral OCR 3 introduces major improvements over OCR 2 with claimed 74% win rate on forms, handwriting, scanned documents, and complex tables. Pricing set at $2 per 1,000 pages ($1 with batch API).
Major release unifying reasoning, multimodal, and coding capabilities. First Mistral model with configurable reasoning effort parameter and native multimodal support in the Small family.
SDK adds support for Claude Opus 4-8 model with mid-conversation system blocks and granular output token usage details.
Music v2 adds mid-track genre switching, section-based composition, and targeted editing capabilities. Released 10 months after Music v1.
Initial release of LocateAnything-3B with Parallel Box Decoding architecture, trained on 12M images across natural scenes, robotics, driving, GUI, and document domains.
Initial release of Gemini Omni Flash, the first tier of Google's multimodal video generation model with avatar cloning and physics modeling capabilities.
Initial release of Stable Audio 3 Medium with 2B parameters, supporting variable-length audio generation up to 6+ minutes with sub-2-second inference times on H200 GPU.
DeepSeek permanently reduced V4 Pro pricing by 75%, dropping input tokens to $0.003625 per million and output tokens to $0.87 per million. Previously promotional pricing is now permanent.
Initial release of diffusion language model family trained on 1.3T pretraining tokens and 45B fine-tuning tokens. Supports autoregressive, diffusion, and self-speculation generation modes with up to 6.4× speedup over traditional AR models.
Hy-MT2 represents a major version release with new 1.8B, 7B, and 30B-A3B model sizes supporting 33 languages, extreme quantization via AngelSlim, and the IFMTBench benchmark for instruction-following evaluation.
Gemini 3.5 Flash now powers Google's AI Mode search, adding support for multimodal inputs including images, video files, and Chrome tabs with improved intent anticipation.
Mistral Medium 3.5 merges instruction-following, reasoning, and coding into a 128B dense model with 256k context. Released as open weights under modified MIT license with configurable reasoning effort and new vision encoder.
Initial release of Command A+ open source model featuring 25B active parameters in a 218B parameter sparse mixture-of-experts architecture with vision, tool use, and reasoning capabilities.
Flagship release of Qwen3.7 series with 1M token context window and agent-first design. Notable improvements in coding and agentic performance over prior Qwen generations with explicit prompt caching support.
Microsoft released Fara1.5-27B, a 27B-parameter vision-only computer use agent fine-tuned from Qwen3.5-27B, supporting 262K context and MIT licensing. The model introduces built-in 'critical points' safety pausing for irreversible actions and is designed for deployment inside Microsoft's MagenticLite sandbox.
Initial release of Grok Build 0.1, xAI's first coding-specialized model with 256K context window designed for agentic workflows and CLI integration.
Gemini 3.5 Flash is Google's first model in the 3.5 series, claiming improved agentic and coding capabilities at reduced cost compared to frontier models. Features enhanced safety measures with reasoning checks before responses.
OlmoEarth v1.1 reduces compute costs by up to 3x through token sequence length reduction, collapsing Sentinel-2 resolution bands into single tokens while maintaining similar benchmark performance to v1.
Initial release of Qianfan-OCR-Fast, a specialized OCR model with 66K context window and domain-specific training for improved document processing performance.
Complete architecture rebuild from XLM-RoBERTa to ModernBERT, expanding context from 512 to 32K tokens (64x increase) and adding code retrieval. The 97M model achieves 60.3 on MTEB Multilingual Retrieval (+12.2 over R1), highest in its size class.
First release of Gemini Omni family supporting multimodal video generation with conversational editing. Speech and audio editing capabilities withheld pending safety testing.
Fast-mode variant of Claude Opus 4.7 with identical capabilities but prioritized output speed at 6x premium pricing ($30/$150 per 1M tokens vs standard rates).
Initial release of Perceptron Mk1, a vision-language model specializing in video understanding and spatial annotation with optional reasoning capabilities.
Initial release of Trinity Large Thinking as a free, open source reasoning model with 262K context window and focus on agentic workloads.
Gemma 4 E4B assistant introduces Multi-Token Prediction architecture for speculative decoding, achieving up to 2x inference speedup. Features 4.5B effective parameters with Per-Layer Embeddings optimized for on-device deployment.
Initial release of Ring-2.6-1T, a 1T parameter thinking model with 63B active parameters, featuring adaptive reasoning and 262K context window optimized for agent workflows.
GPT-5.5-Cyber is a variant of GPT-5.5 with relaxed safeguards specifically for vetted cybersecurity teams. The model is trained to be more permissive on security-related tasks compared to the standard GPT-5.5 release.
First OpenAI voice model with GPT-5-class reasoning capabilities, designed for live voice interactions with tool calling and natural conversation flow.
Google launches Flash Lite variant of Gemini 3.1 at 50% cost reduction with maintained 1M context window and four-level reasoning system.
Initial release of Multi-Token Prediction assistant model for Gemma 4 26B A4B, enabling up to 2x inference speedup through speculative decoding while maintaining identical output quality.
Major release introducing Gemma 4 family with four model sizes (E2B, E4B, 26B A4B MoE, 31B dense), Multi-Token Prediction drafters for 2x speedup, extended context windows up to 256K, enhanced multimodal capabilities, and improved reasoning performance.
Initial release of CoBuddy code generation model with 131K context window, native tool calling, and reasoning support, available for free on OpenRouter.
GPT-5.5 Instant replaced GPT-5.3 Instant as the default ChatGPT model with significant accuracy improvements and more concise response formatting. OpenAI claims 52.5% fewer hallucinations on high-stakes prompts and reduced use of emojis and unnecessary formatting.
Granite Speech 4.1 2B introduces dual-head CTC encoder, frame importance sampling, improved multilingual ASR accuracy, and punctuation/truecasing across all languages. Two new variants add speaker attribution with timestamps and non-autoregressive architecture.
IBM released Granite 4.1 family in 3B, 8B, and 30B sizes under Apache 2.0 license. Unsloth released 21 GGUF quantized variants of the 3B model.
April 2026
Granite 4.1 8B is a new release in IBM's Granite 4.1 family, offering an 8B-parameter dense decoder-only model with enterprise-focused capabilities. Released under Apache 2.0 license with 131K context window and multilingual support.
R2 upgrades architecture from XLM-RoBERTa to ModernBERT, extends context from 512 to 32,768 tokens, expands vocabulary to 262K tokens, and adds Matryoshka dimension reduction support. Performance improves by 11.8 points on MTEB Retrieval.
Granite 4.1 30B introduces enhanced tool-calling, improved instruction following through updated SFT-RL pipeline, and 131K context window. Released under Apache 2.0 license with competitive performance on code and reasoning benchmarks.
Initial release of Nemotron 3 Nano Omni, a 30B-parameter multimodal MoE model designed as a perception sub-agent for enterprise systems. Features hybrid Transformer-Mamba architecture with specialized video processing and extended reasoning capabilities.
Initial release of Nemotron 3 Nano Omni, a 31B-parameter MoE multimodal model with video, audio, image, and text understanding, 256K context window, and dedicated reasoning mode with chain-of-thought capabilities.
Initial release of 13-billion-parameter vintage language model trained exclusively on public domain texts published before 1931.
Initial release of Nemotron 3 Nano Omni, a multimodal MoE model with 30B total parameters (3B active) combining video, audio, image, and text understanding in a single inference pass with 131K token context.
Second-generation XS size model with enhanced tool calling and reasoning capabilities, quantized to fp8 for production efficiency.
Initial release of Nemotron-3-Nano-Omni-30B-A3B, a multimodal MoE model with 31B parameters combining video, audio, image, and text understanding with reasoning capabilities.
First omni-modal release in Nemotron 3 line, adding audio and video capabilities to previous vision-language model. Uses new 30B-A3B MoE architecture with hybrid Mamba-Transformer design.
Initial release of Owl Alpha, OpenRouter's first foundation model designed specifically for agentic workloads with native tool use and 1M+ context window.
Laguna XS.2 introduces a 33B parameter MoE architecture with 3B activated parameters per token, featuring mixed sliding window and global attention layers, native reasoning support, and optimization for local deployment.
Initial release of GPT Mini Latest with 400,000 token context window and auto-redirect to newest GPT Mini family version.
Qwen3.6 Flash introduces multimodal support for text, images, and video with a 1M token context window. Features tiered pricing structure with base rates of $0.25/$1.50 per 1M tokens for prompts under 256K tokens.
Google released Gemini Flash Latest as a dynamic router that automatically redirects to the newest Gemini Flash model, featuring 1,048,576 token context and reasoning capabilities.
Moonshot AI introduced a router endpoint that automatically redirects to the most current model in the Kimi family, featuring a 262,144 token context window.
Updated Qwen3.5 Plus with 1M token context window and tiered pricing above 256K tokens.
Initial release of Qwen3.6 Max Preview, a proprietary 1 trillion parameter sparse MoE model with 262K context window and integrated thinking mode for agentic workflows.
MiMo-V2.5-Pro introduces 1.02T total parameters (up from 310B in V2.5) with 42B active parameters, extends context to 1M tokens, and adds hybrid attention architecture with sliding window and global attention patterns. The model achieves 99.6% on GSM8K and maintains coherence at extreme context lengths.
Google released a dynamic router endpoint that automatically redirects to the latest Gemini Pro model, featuring 1,048,576-token context and reasoning capabilities.
New sparse MoE model with 35B total parameters but only 3B active per token, featuring 262K context window, multimodal support, and integrated reasoning mode.
Qwen3.6 27B introduces video processing capabilities alongside existing text and image support, with a 262K context window and built-in thinking mode for agentic coding and reasoning tasks.
GPT-5.5 Pro introduces a 1M+ token context window optimized for complex reasoning tasks, agentic coding, and multi-step workflows with multimodal support.
DeepSeek-V4-Flash is a new 284B-parameter MoE model with 13B activated parameters, featuring hybrid attention architecture that reduces inference costs by 73% at million-token context lengths. Introduces three reasoning effort modes and achieves competitive performance with frontier models on coding and mathematical reasoning.
Major release with 1.6T total parameters, 1M token context window, and substantial efficiency improvements over V3.2. DeepSeek claims near-frontier performance at a fraction of the cost.
Major version release with claimed competitive performance against leading US models. DeepSeek emphasizes improved coding capabilities and domestic chip compatibility.
Initial release of Ling-2.6-1T, a 1 trillion parameter instruct model with 262K context window and fast-thinking architecture designed for cost-efficient agent deployments.
Qwen3.6-27B is a 27B dense model that claims flagship-level coding performance surpassing the 397B Qwen3.5-397B-A17B. Available at 55.6GB full size or 16.8GB quantized.
Gemma 4 E2B demonstrated running as a vision-language agent on NVIDIA Jetson Orin Nano Super (8GB), autonomously deciding when to access webcam based on conversational context with no hardcoded triggers.
Initial release of Privacy Filter, a 1.5B-parameter bidirectional token classifier for detecting 8 PII categories with 128K context window and Apache 2.0 license.
Initial preview release of Hy3, Tencent's Mixture-of-Experts model with configurable reasoning modes designed for production agentic workflows.
Initial release of Trinity Large Preview, a 400B-parameter sparse MoE model with 13B active parameters per token, featuring up to 512K context window support and optimization for agentic workflows.
Initial release of Pareto Code Router with dynamic model selection based on min_coding_score parameter (0-1 scale) and 200K context window.
Major update to OpenAI's image generation capabilities with significantly improved text rendering and UI element generation accuracy.
GPT-5.4 Image 2 combines GPT-5.4 reasoning capabilities with image generation from GPT Image 2, supporting text, image, and file inputs with a 272K token context window.
Major update to OpenAI's image generation model with significantly improved quality, maximum resolution of 3840x2160 pixels, and pricing at $30 per million output tokens.
Initial release of Ling-2.6-flash with 104B total parameters (7.4B active), featuring 262K context window and optimized for agent applications with fast response times.
First release of Kimi K2.6 introduces 1T-parameter MoE architecture with agent swarm capabilities, 256K context, and competitive performance on SWE-Bench and agentic benchmarks.
Initial release of Qianfan-OCR-Fast, a specialized OCR model with 65K context window and free pricing. Claims performance improvements over Qianfan-OCR.
GR00T N1.7 upgrades to Cosmos-Reason2-2B VLM backbone and adds EgoScale pre-training on 20,854 hours of human egocentric video, improving dexterity and generalization over N1.6.
Python SDK v0.96.0 adds Claude Opus 4-7 support, introduces token budgets for cost management, and user profiles for personalized interactions.
Opus 4.7 introduces significantly more aggressive safety guardrails that automatically detect and block requests related to cybersecurity uses, resulting in a 10-15x increase in reported false positive refusals.
Major release introducing multi-modal 3D world generation capabilities, replacing video-based outputs with persistent 3D assets. Centered on WorldMirror 2.0, a 1.2B parameter model for reconstruction.
Gemini 3.1 Flash TTS introduces audio tags for granular control over vocal style, pace, and delivery through natural language commands. The model achieves an Elo score of 1,211 and includes mandatory SynthID watermarking.
Fine-tuned variant of GPT-5.4 specifically built for defensive cybersecurity work with reduced restrictions on security-related tasks. Initial access limited to verified security professionals through Trusted Access for Cyber program.
Initial preview release of Gemini 3.1 Flash TTS, Google's first prompt-controlled text-to-speech model that accepts theatrical-style direction for voice characteristics, accents, and delivery style.
Anthropic released Mythos, a security-focused AI model designed to identify and exploit zero-day vulnerabilities. The model remains unreleased publicly due to security concerns.
Initial release of Trinity-Large-Thinking, a 400B parameter open-weight reasoning model with 256 mixture-of-experts (13B active per token) optimized for agent tasks. Apache 2.0 licensed, trained on 17 trillion tokens over 33 days on 2,048 Nvidia B300 GPUs.
Refreshed version of LFM2-VL-450M with updated LFM2.5-350M backbone. Adds bounding box prediction and function calling capabilities while improving performance across vision and language benchmarks.
Released HY-Embodied-0.5 suite with MoT-2B (2.2B active parameters, 4B total) and 32B variants. MoT-2B trained on 200B+ tokens of embodied data outperforms similarly-sized competitors on 22 embodied benchmarks.
Amazon Bedrock now enables supervised fine-tuning, reinforcement fine-tuning, and model distillation for Nova 2 Lite. Fine-tuned models deploy on-demand at standard inference pricing without provisioned capacity.
NVIDIA released Alpamayo 2 Super, a new 34B-parameter vision-language-action model combining a 32B VLM backbone with a 2.3B diffusion-based action decoder for autonomous vehicle perception, planning, and reasoning tasks.
Claude Mythos Preview demonstrates unprecedented capability in autonomous vulnerability discovery and exploitation, finding decades-old bugs in OpenBSD, FFmpeg, and FreeBSD. Deployment restricted to 11-partner coalition for defensive cybersecurity use only.
Mythos is Anthropic's new frontier model, positioned as larger and more intelligent than its Opus models. It is deployed exclusively through Project Glasswing for defensive cybersecurity work with 40+ vetted partner organizations, with claims of identifying thousands of zero-day vulnerabilities during early testing.
Amazon Nova 2 Sonic enables real-time conversational podcast generation with 1M token context window and native support for seven languages through Amazon Bedrock.
Initial release of Harrier embedding model. Trained on 2B+ examples with GPT-5 synthetic data. Achieves top ranking on MTEB v2 multilingual benchmark with 131K context window.
Claude Mythos is Anthropic's specialized model for cybersecurity vulnerability discovery, designed to identify critical flaws in operating systems, browsers, and software. The model shows improvements over Claude Opus 4.6 in reasoning, agent-based capabilities, and coding.
Claude Mythos Preview released under restricted access through Project Glasswing. Model demonstrates exceptional cybersecurity research capabilities including discovery of 27-year-old OpenBSD TCP SACK vulnerability and Linux privilege escalation flaws.
Initial release of Bonsai 8B 1-bit quantized model. Achieves 14x compression with claimed competitive performance on standard benchmarks. Also released Bonsai 4B and Bonsai 1.7B variants.
Gemma 4 E4B adds multimodal capabilities (text, image, audio), extended 128K context window, native reasoning modes, and function-calling support compared to Gemma 3. Achieves 69.4% MMLU Pro with 4.5B effective parameters optimized for mobile and edge deployment.
Gemma 4 introduces multimodal support (text, image, video, audio on small models), extended context windows (128K-256K tokens), configurable reasoning modes, and native function calling. Available in four sizes with both dense and MoE architectures.
Zhipu AI released GLM-5V-Turbo, adding multimodal capabilities to its GLM-5 series. The model generates code from design mockups and video inputs while maintaining text-only coding performance, integrating directly with Claude Code and OpenClaw agents.
Tencent released OmniWeaving, an open-source unified video generation model with reasoning capabilities and compositional video creation. Built on HunyuanVideo-1.5, it supports eight video generation tasks and introduces IntelligentVBench benchmark.
Gemma 4 26B A4B uses Mixture-of-Experts with 3.8B active parameters for efficient inference. Features 256K context window, multimodal input (text/image), native reasoning modes, and function-calling for agentic workflows.
Microsoft released MAI-Transcribe-1, a speech-to-text model achieving lowest FLEURS benchmark word error rate at 2.5x faster inference than Azure Fast. Priced at $0.36 per audio hour, supporting 25 languages and challenging recording conditions.
Google DeepMind introduces Gemma 4 31B with multimodal input (text and images), 256K context window, configurable reasoning mode, and native function calling. Free release under Apache 2.0 license.
Gemma 4 introduces multimodal support, 256K context window, Apache 2.0 permissive licensing, and mixture of experts variant. First major version update with explicit focus on enterprise deployment without data usage restrictions.
Gemma 4 introduces multimodal capabilities with native image and audio support, extended 128K context window, built-in reasoning modes with configurable thinking, and hybrid attention architecture combining local and global attention for efficiency.
Gemma 4 introduces multimodal capabilities (text, image, video support), reasoning modes, 256K context windows, and Mixture-of-Experts architecture. The 26B A4B variant uses sparse activation for near-dense-31B performance with 4B-model inference speed.
Qwen3.6 Plus introduces hybrid linear attention with sparse mixture-of-experts routing, achieving 78.8 on SWE-bench Verified. Major improvements in coding, reasoning, and multimodal capabilities over 3.5 series.
Gemma 4 introduces four model sizes (2B-31B) with improved reasoning and agentic capabilities. Apache 2.0 licensing replaces previous restrictions. 31B model ranks #3 on Arena AI leaderboard.
NVIDIA released NVFP4-quantized version of Google DeepMind's Gemma 4 31B IT model optimized for consumer GPU inference. Maintains 256K context window and multimodal capabilities with <0.5% performance degradation on reasoning and coding benchmarks.
Qwen 3.6 Plus introduces a hybrid architecture with linear attention and sparse mixture-of-experts routing, delivering major improvements in agentic coding, front-end development, and reasoning over the 3.5 series. Achieves 78.8 on SWE-bench Verified.
Initial release of Falcon Perception 0.6B early-fusion Transformer for open-vocabulary grounding and segmentation. Introduces Chain-of-Perception output interface and PBench diagnostic benchmark with five capability levels.
Holo3-122B-A10B released with 78.85% OSWorld score using mixture-of-experts architecture (122B total, 10B active parameters). Trained via agentic learning flywheel with synthetic data augmentation and curated reinforcement learning. Holo3-35B-A3B variant open-sourced under Apache 2.0.
Meta's first closed-source AI model with planned paid developer access, marking a strategic shift from the open-source Llama series. Released from Meta Superintelligence Labs under new AI leadership.
March 2026
Google released Veo 3.1 Lite, a cost-optimized video generation model priced at less than 50% of Veo 3.1 Fast. Designed for high-volume applications with same generation speed as Veo 3.1 Fast.
Grok 4.20 Multi-Agent is a specialized variant designed for collaborative agent-based workflows with parallel agent coordination. Scales agent count based on reasoning effort: 4 agents at low/medium effort, 16 agents at high/xhigh effort.
Microsoft released the Harrier-OSS embedding model family with three parameter sizes (270M, 600M, 27B) supporting multilingual inputs, 32K token context, and knowledge distillation techniques. The 27B variant achieves 74.3 on MTEB v2 benchmark.
Granite 4.0 3B Vision introduces a compact vision-language model optimized for enterprise document processing with DeepStack Injection architecture and ChartNet dataset training. Shipped as LoRA adapter on Granite 4.0 Micro for modular text-only fallback support.
Qwen3.5-Omni expands from Qwen3-Omni with 8x context window increase (32K to 256K tokens), 6x language support expansion (11 to 74 languages), hybrid attention-MoE architecture, and ARIA token interleaving for improved real-time speech synthesis. Demonstrates emergent code-generation capability from spoken and video input.
Google releases Lyria 3 Pro Preview, a music generation model producing full-length songs with vocals, lyrics, and instrumental arrangements. Priced at $0.08 per song through the Gemini API with 1M token context window.
Lyria 3 Clip Preview introduces Google's music generation model to the Gemini API with clip-based pricing at $0.04 per 30-second generation.
Microsoft releases Harrier-OSS-v1 family of multilingual embedding models in three sizes (270M, 0.6B, 27B parameters) trained with contrastive learning and knowledge distillation. The 0.6B variant achieves 69.0 MTEB v2 score with 32,768 token context window and supports 45+ languages.
Qwen 3.6 Plus Preview introduces a hybrid architecture with 1M token context window, improved reasoning and agentic behavior over 3.5 series. Available free on OpenRouter with data collection for model improvement.
KAT-Coder-Pro V2 builds on earlier KAT-Coder versions with enhanced agentic coding capabilities for large-scale production environments and added web design generation for landing pages and presentation decks.
Cohere releases Transcribe, a 2B parameter open-source speech recognition model with 5.42% WER on the Hugging Face leaderboard, supporting 14 languages under Apache 2.0 license.
Google released Gemini 3.1 Flash Live, an audio-focused model optimized for multilingual conversations. The model powers the expansion of Search Live to 200+ countries with claimed improvements to response speed and conversation naturalness.
NVIDIA released gpt-oss-puzzle-88B, an inference-optimized 88B-parameter mixture-of-experts model derived from gpt-oss-120B using the Puzzle NAS framework. Achieves 1.63× throughput improvement on long-context and up to 2.82× on single H100s while maintaining parent accuracy through heterogeneous expert pruning, selective window attention, and knowledge distillation with RL optimization.
Dreamina Seedance 2.0 launches in CapCut with IP safeguards including face-detection blocks, unauthorized IP generation prevention, and invisible watermarking to identify AI-generated content.
Gemini 3.1 Flash Live improves upon 2.5 Flash Native Audio with enhanced acoustic recognition, background noise filtering, lower latency, and extended conversation context.
Apple developed RubiCap, a rubric-guided reinforcement learning framework for dense image captioning that achieves state-of-the-art results with 2B-7B parameter models, outperforming competitors up to 72B parameters.
Initial release of MolmoWeb with 4B and 8B parameter variants. Includes full training dataset (MolmoWebMix), model weights, and evaluation tools under Apache 2.0 license.
Expanded from 30-second to 3-minute song generation with improved structural composition control. Added support for specifying discrete song elements.
Nemotron 3 Super now available on Amazon Bedrock as fully managed serverless inference. 120B parameter MoE model with 12B active parameters, 256K context, claims 5x throughput improvement and 2x accuracy gain over previous version.
Xiaomi released MiMo-V2-Pro as 3x larger successor to MiMo-V2-Flash (Dec 2025), reaching 1T parameters with 42B active per request. Benchmarks place it 3rd globally on PinchBench/ClawEval, nearly matching Claude Opus 4.6 on coding (78% vs 80.8%) while costing 80% less per input token.
Composer 2 launched with frontier-level coding intelligence, built on Moonshot AI's open-source Kimi 2.5 model with additional reinforcement learning training applied by Cursor (75% of final compute).
M2.7 introduces autonomous participation in its own development through 100+ self-optimization rounds, achieving 30% performance improvement on internal coding tasks and competitive benchmark scores against leading Western models.
MAI-Image-2 improves upon MAI-Image-1 with enhanced photorealism, natural lighting, and notably adds reliable text rendering capabilities for practical design applications. The model ranks third on Arena.ai leaderboard, up from ninth place for the previous version.
Composer 2 achieves 61.3 on CursorBench (+38% vs Composer 1.5) through improved pretraining and reinforcement learning on long-horizon tasks. Pricing set at $0.50/$2.50 per 1M tokens, undercutting Claude and GPT-4 by 60-90%.
MiMo-V2-Pro is Xiaomi's flagship foundation model launch featuring 1T+ parameters and 1M context window optimized for agent systems and complex workflow orchestration.
Xiaomi launches MiMo-V2-Pro, their flagship 1T-parameter foundation model with 1M context. Optimized for agentic scenarios, ranking among global top tier on standard benchmarks.
Xiaomi debuts MiMo-V2-Omni, a frontier omni-modal model processing image, video, and audio natively. 262K context with strong agentic capabilities including visual grounding and code execution.
Qwen3.5-Max-Preview launches as Alibaba's largest model (1T+ params) in the Qwen3.5 series. 262K context with thinking mode (82K CoT). Beats previous flagship on reasoning, multilingual, and agentic tasks. $1.20/$6.00 per 1M tokens.
GPT-5.4 mini introduces major improvements in coding, reasoning, and computer control capabilities over GPT-5 mini. Model runs over 2x faster and achieves near-full-GPT-5.4 performance on multiple benchmarks while consuming 30% of quota in agentic systems.
GPT-5.4 Nano launches as the smallest, fastest member of the GPT-5.4 family. 400K context, multimodal input, optimized for high-volume agentic tasks at $0.20/$1.25 per 1M tokens.
GPT-5.4 mini, OpenAI's fastest variant of GPT-5.4, is now generally available in GitHub Copilot. The model claims to be the highest-performing mini offering for coding tasks.
Initial release of MiniMax M2.7 — next-gen LLM with multi-agent collaboration, 204K context, SWE-Pro 56.2%, Terminal Bench 2 57.0%
Initial release of Nemotron-3-Nano-4B-GGUF, a quantized (Q4_K_M) 4B parameter edge model with hybrid Mamba-2 architecture. Supports controllable reasoning modes and 262K context window for edge AI applications including gaming NPCs and local voice assistants.
Mistral Small 4 launches unifying Magistral reasoning, Pixtral multimodal, and agentic coding into one model. 262K context at $0.15/$0.60 per 1M tokens.
ByteDance releases Seedream 4.5, their latest image generation model with major quality improvements over Seedream 4.0.
GLM 5 Turbo launches with 203K context and fast inference optimized for agent-driven environments. Improved complex reasoning over base GLM 5 at $0.96/$3.20 per 1M tokens.
Minimax released M1 80k, expanding its M1 model family with an 80,000-token context window for extended document processing.
Nvidia released Llama 3.1 Nemotron 70B Instruct, an instruction-tuned variant of Meta's Llama 3.1 70B model optimized for developer applications.
Mistral AI releases Pixtral Large, a new multimodal model supporting image and text inputs with 128K context window.
Minimax releases M1 40k model with 40,000-token context window. Initial release with limited publicly disclosed specifications.
Google launched Ask Maps, integrating Gemini AI into Google Maps to allow users to ask complex contextual navigation questions. The chatbot personalizes responses based on user search history and saved locations.
Nvidia releases Nemotron 3 Super, a 120B hybrid MoE model with 1M context window, latent expert routing, and multi-token prediction. Fully open-weight under NVIDIA Open License.
ByteDance launches Seed-2.0-Lite, a cost-efficient multimodal enterprise model with 262K context. Strong agent capabilities at $0.25/$2.00 per 1M tokens.
NVIDIA releases Nemotron-3-Super-120B-A12B-BF16, a 120 billion parameter model with latent MoE architecture for efficient text generation across 8 languages.
NVIDIA releases Nemotron-3-Super-120B, a 120B parameter model with latent MoE architecture optimized for conversational tasks across 8 languages.
StepFun releases Step-3.5-Flash-Base as an open-source text generation model optimized for efficient inference under Apache 2.0 license.
Initial release of Qwen3.5-2B, a 2-billion-parameter multimodal model supporting image and text processing.
Qwen3.5-4B released as a 4 billion parameter multimodal model supporting image and text inputs. Apache 2.0 licensed for open-source use.
Gemini 3.1 Flash Lite Preview launches as Google's high-efficiency model for high-volume use. 1M context at $0.25/$1.50, outperforms Gemini 2.5 Flash Lite.
Qwen3.5-9B released as multimodal 9-billion parameter model supporting image and text inputs. Available under Apache 2.0 license on Hugging Face.
Qwen3.5-0.8B released as an 800-million-parameter multimodal model for edge inference. Supports image and text inputs under Apache 2.0 licensing.
Initial release of Context-1, a 20B parameter Mixture of Experts retrieval agent model trained for multi-hop search with self-editing context capabilities.
Released FP8-quantized version of Qwen3.5-35B-A3B, reducing memory requirements while maintaining multimodal capabilities. Compatible with Transformers endpoints and Azure deployment.
Gemini 3.1 Flash-Lite achieves 2.5x faster first-token latency than Gemini 2.5 Flash with 360 tokens/second throughput. Output pricing increased to $1.50 per million tokens from $0.40.
February 2026
Qwen3.5-35B-A3B-Base released as a 35-billion parameter multimodal model with Apache 2.0 license. Part of the Qwen3.5 mixture-of-experts family.
Nano Banana 2 (Gemini 3.1 Flash Image Preview) debuts as Google's fastest image generation model. Pro-level quality at Flash speed with $0.50/$3.00 per 1M tokens.
Seed-2.0-Mini launches targeting latency-sensitive scenarios with 262K context and four reasoning effort levels at $0.10/$0.40 per 1M tokens.
Qwen3.5-Flash debuts with 1M context and ultra-low $0.065/$0.26 pricing. Hybrid architecture delivers a leap in inference efficiency over Qwen 3 series.
Qwen3.5-27B released as a 27-billion parameter multimodal model supporting image-text-to-text tasks. Available under Apache 2.0 license with transformer endpoint compatibility.
Qwen3.5-35B-A3B released as open-weight multimodal model with 35B parameters. Apache 2.0 licensed, supports image and text inputs with conversational capabilities.
Mercury 2 launches. The first reasoning-capable diffusion LLM. Hits 1,009 tokens/sec on NVIDIA Blackwell with 1.7s end-to-end latency — 5x faster than speed-optimized autoregressive models. Tunable reasoning, native tool use, schema-aligned JSON output. OpenAI-compatible API.
Cohere Labs releases tiny-aya-global, a multilingual text generation model fine-tuned from tiny-aya-base to support conversational tasks across 100+ languages including major and low-resource languages.
Gemini 3.1 Pro enters public preview in GitHub Copilot with focus on efficient edit-then-test loops and agentic coding capabilities.
Lyria 3 integrated into Gemini, enabling 30-second music track generation with vocals, lyrics, and cover art from text prompts or uploaded media.
Alibaba released Qwen 3.5 series claiming performance parity with proprietary frontier models while optimized for commodity hardware, directly challenging closed-source AI model economics.
Claude Opus 4.6 — major GPQA and reasoning improvements; ARC-AGI-2 jump from 37.6% to 68.8%.
GLM-5.1 introduces sustained agentic reasoning over hundreds of iterations with improved performance on SWE-Bench Pro (58.4%, +3.3pp vs. GLM-5) and NL2Repo (42.7%, +6.8pp vs. GLM-5). The model maintains productivity across longer problem-solving sessions through iterative experimentation and strategy revision.
Qwen3.5 397B A17B launches as a hybrid linear-attention + sparse MoE vision-language model. 262K context at $0.39/$2.34 per 1M tokens with state-of-the-art performance.
Qwen3.5 Plus launches with 1M context and strong cross-domain performance in academia, finance, marketing, programming, and science at $0.30/$1.20 per 1M tokens.
Gemini 3.1 Pro released as upgrade to Gemini 3 Pro. Enhanced reasoning for complex multi-step problems. Preview access via AI Studio.
MiniMax-M2.5 launches.
Claude Sonnet 4.6 brings improved multi-step reasoning and stronger agentic task performance over 4.5, at the same price.
GPT-5.3-Codex launches. First model combining the Codex and GPT-5 training stacks — best-in-class code generation plus general-purpose reasoning in one model, ~25% faster than its predecessor.
Step-3.5-Flash launches.
Grok 4 mini brings next-generation reasoning at low cost. Significantly outperforms Grok 3 mini on AIME and competitive coding tasks.
January 2026
Kimi K2.5 launches as Moonshot AI's native multimodal model with state-of-the-art visual coding and agent swarm paradigm. 262K context at $0.45/$2.20 per 1M tokens.
Claude Sonnet 4.5 January patch with stability improvements for long agentic sessions and better tool use across multi-turn workflows.
Gemini 2.5 Pro preview January update with longer stable thinking output windows and improved factual grounding on complex queries.
Global rollout suspended following copyright cease-and-desist letters from Disney and Paramount Skydance over use of copyrighted training material.
December 2025
Gemini 3 Flash launches. Fast, cost-efficient Gemini 3 workhorse with near-Pro reasoning at a fraction of the cost. 1M context. Rolled out globally in the Gemini app, Search AI Mode, and the API on day one.
DeepSeek V3 December update with improved instruction following and expanded Chinese-English code-switching performance.
Mistral Small 3.1 December update with improved vision accuracy and expanded support for structured data extraction from images.
Gemini 3.0 Pro debuts as first Gemini 3 generation model. Upgraded reasoning and native multimodal understanding. Quickly superseded by 3.1.
November 2025
FLUX.2 launches. Second-generation FLUX with 4MP photoreal output, multi-reference character consistency, accurate text rendering, and editing in one checkpoint. Open-weights dev variant alongside pro/flex API tiers.