AI Model Changelog
Every version update across all major AI models.
August 2026
DeepSeek released an experimental vision-enabled variant of V4 Flash 0731, adding image understanding capabilities while matching the base model's text performance. The sparse MoE model uses 13B active parameters out of 284B total with a 1M token context window.
GEN-1.5 introduces in-context learning for robot manipulation, allowing a single short demo to be loaded as a 'physical prompt' instead of requiring dataset collection and training. Generalist reports 59% zero-shot success and 83% after brief fine-tuning across ten test tasks, though results are unverified.
Ox Alpha is a new stealth reasoning model released anonymously through OpenRouter, offering a free preview with a 1 million token context window for coding and agentic use cases.
Tencent released Hy-MT2-30B-A3B, a mixture-of-experts translation model with 30B total and 3B active parameters, supporting 33 language pairs and 5 Chinese dialect/minority-language pairs across five distinct translation workflows.
Qwen3.8-27B is a 27B-parameter dense vision-language model built on the Qwen3.5 architecture, improving on Qwen3.6-27B and Qwen3.7-Plus across coding, agentic, and multimodal benchmarks. It adds native image/video understanding, flexible thinking control, and a 262,144-token native context window extensible to 1 million tokens.
Qwen3.8-27B is a 27B dense vision-language model built on the Qwen3.5 architecture, offering native 262K context extensible to 1M tokens and configurable reasoning depth. This release provides FP8-quantized weights for efficient deployment across standard inference frameworks.
Google released Gemini 3.7 Flash, a multimodal model with a 1,049K token context window aimed at fast agentic workflows and coding, priced at $0.38/M input and $1.88/M output tokens.
MiniMax released Music 3, a new open-weight text-to-music model that generates complete, structurally coherent songs up to five minutes long from lyrics and detailed music descriptions.
Writer released Palmyra X6, a post-trained variant of Z.ai's open source GLM-5.2, paired with an upgraded agentic harness, claiming up to 50% lower token costs for basic tasks.
Alibaba released Qwen3.8-2.4T-A95B-FP8, an open-weight, FP8-quantized MoE checkpoint with 2.4T total and 95B activated parameters, serving as the foundation for the hosted Qwen3.8-Max API. The model extends context to 1M tokens and shows significant benchmark gains over Qwen3.7-Max on coding and agentic tasks, according to Alibaba.
Initial public release of SL2T, Google DeepMind's massively multilingual sign-language-to-text model, deployed in Gboard and Live Transcribe on Pixel 11 for ASL-to-English translation.
MAI Code 1.1 Flash improves on MAI-Code-1-Flash with a 25% token-efficiency gain and lower pricing, but trails DeepSeek-V4-Flash-0731 on Terminal Bench 2.1 and costs substantially more per output token.
Alibaba released Qwen3.8-2.4T-A95B, a 2.4-trillion-parameter MoE model with 95B active parameters, built on the Qwen3.5 architecture with hybrid Gated DeltaNet and Gated Attention layers. It introduces reasoning_effort and preserve_thinking controls and claims Qwen-Max-class performance in an open-weight release for the first time.
LFM2.5-VL-3B upgrades Liquid AI's edge vision-language model with stronger screen/UI understanding, grounding, multi-image reasoning, and significantly improved function calling, while maintaining a 3.1B-parameter footprint for on-device deployment.
xAI released Grok 4.6 with a 500K token context window and multimodal input support (text, image, file), now available via OpenRouter's API. The company claims it is their smartest model yet, with frontier performance on coding, knowledge work, and STEM tasks, though no independent benchmark scores were disclosed.
NVIDIA released a new NVFP4-quantized checkpoint of Nemotron 3.5 Lightning, a 30B-total/3B-active hybrid Mamba-2/MoE/Attention model with 1M-token context, optimized for single-GPU deployment on DGX Spark or H100 hardware. The release includes post-training quantization and speculative decoding support to preserve BF16-level accuracy at lower inference cost.
Nemotron 3.5 Lightning succeeds Nemotron 3 Nano 30B A3B, gaining nine points on the Intelligence Index (15 to 24) and large jumps on agentic benchmarks (GDPval-AA v2 Elo up 334 points, Terminal-Bench v2.1 up from 7% to 24.3%) while becoming the fastest model in its comparison class at nearly 670 tokens per second.
Sakana AI released Namazu, a reasoning model fine-tuned from Kimi K2.6 for Japanese language and business contexts, with a 262K token context window.
Upstage released Solar Pro 4, a text-to-text model with a 524K token context window aimed at agentic workflows, document processing, and coding, priced at $0.03/M input and $0.12/M output tokens.
ByteDance Seed released Seed-2.0-Code, a new coding-focused model with a 262K token context window and multimodal input support, designed for agentic coding and frontend development tasks.
NVIDIA released the full-precision BF16 reference weights for Nemotron 3.5 Lightning, a 30B-parameter hybrid Mamba-MoE-Attention model with 3B active parameters and up to 1M token context. This release is intended for fine-tuning and quantization rather than direct deployment, with a companion NVFP4 checkpoint available for optimized inference.
New 29.6B-parameter open-weight model with a built-in perception encoder, distilled from Muse Spark, designed for on-device agentic use with 4-bit quantization and speculative decoding.
Magpie TTS Multilingual adds Modern Standard Arabic, Korean, and Brazilian Portuguese, expanding coverage to 12 languages, and improves speech quality (lower CER, higher SSIM) on several existing languages including French and Spanish.
Meta's first open-weight model since Llama 4 (spring 2025), a 30B-parameter agent model released under Apache 2.0, distilled from the larger Muse Spark model and optimized for local, on-device execution.
ByteDance Seed released Seed 2.1 Turbo, a multimodal model with a 262K token context window aimed at coding and agentic software delivery tasks. It is priced at $0.50 per 1M input tokens and $2.50 per 1M output tokens.
GPT-5.6-Cyber is a new specialized variant of GPT-5.6 Sol trained for offensive security research, available exclusively through OpenAI's Daybreak Red access tier. It significantly outperforms its predecessor GPT-5.5-Cyber and standard safety-tuned models on sensitive security query completion and real-world exploit development.
xAI launched Imagine Image 2.0 as a new Quality Mode with editing tools like Magic Wand, Multi-Ref Editing, and Smart Resize. The model ranks second globally on both the Image Edit and Text-to-Image Arena leaderboards, behind OpenAI's GPT-Image-2.
FLUX 3 Video moves from limited access to general availability via the BFL API and select partners, adding native audio, multi-scene generation, and 14+ language lip-sync. BFL claims the model outperforms Seedance 2.0, Gemini Omni Flash, and Minimax H3 on its internal Elo benchmarks.
Meta released Muse Spark 1.2 alongside its first coding agent, Muse Code, with the two trained together for improved coding performance over the standalone Muse Spark 1.1.
Initial release of Shieldstral-1.0-3B, a 3B-parameter policy-adaptive safety classifier built on Ministral-3-3B-Base-2512 with a Pixtral vision encoder. The model replaces fixed moderation categories with natural-language policy queries evaluated at inference time.
Liquid AI released LFM2.5-2.6B, a 2.6B-parameter open-weight model trained via supervised fine-tuning, teacher specialization, on-policy distillation, and agentic RL for local tool-use agents. The model claims competitive performance with models up to 4x its size on instruction-following, tool-use, and agentic benchmarks while running at 220 tok/s on Apple M5 Max hardware.
Mistral released Shieldstral, a 3B open-weights multimodal safety classifier under Apache 2.0 that reframes moderation as policy-adaptive question answering, claiming performance matching guard models up to 7x its size.
K-EXAONE 2.0 scales up from its predecessor's 236B/23B parameter MoE architecture to 750B/37B parameters via upcycling, adding long-context retrieval strength, expanded multilingual coverage (6 to 10 languages), and improved safety benchmark performance.
MiniMax released open weights for H3, a 33B-parameter video model that ranks first in Video Editing and second in Text-to-Video on Artificial Analysis. The 2K resolution module and H3-Context-IR remain closed, capping local generation at 768p.
Initial v1.0 release of NemotronLabs VoiceChat 11B, NVIDIA's first open full-duplex speech model combining a Fast Conformer encoder, Nemotron Nano V2 9B LLM backbone, and TTS decoder into one architecture. The model adds live tool-calling support during natural conversation, a first among open full-duplex voice models according to NVIDIA.
Qwen3.8-Max scales to 2.4 trillion total parameters (95 billion active) and is built on the Qwen3.5 architecture, with training focused on sustaining autonomous, multi-day agentic tasks rather than single-turn responses. It is the first Qwen-Max model to be open-weighted, with weights arriving on Hugging Face and ModelScope next week.
Seedance 2.5 extends maximum clip length to 30 seconds (up from shorter limits in 2.0) and expands reference input support to 30 images, 10 videos, and 10 audio files for multi-character scene construction. ByteDance also claims improved visual fidelity in textures, lighting, and skin detail.
DeepSeek introduced a rolling 'latest' alias for its V4 Flash model family, currently offering a 1,049K token context window at $0.09/M input and $0.18/M output tokens.
July 2026
DeepSeek released V4 Flash 0731, a re-post-trained sparse MoE checkpoint with 13B active/284B total parameters and a 1049K token context window, priced at $0.14/M input and $0.28/M output.
Thinking Machines Lab released Inkling Small, an open-weight multimodal MoE model with 12B active parameters out of 276B total and a 524K context window.
Gemini Robotics ER 2 upgrades from ER 1.6 with continuous video-based progress tracking, faster moment-finding, multi-robot collaboration, and improved safety behavior around humans. It is now available via the Gemini API and Google AI Studio, with private preview on the Gemini Enterprise Agent Platform.
OpenAI cut GPT-5.6 Terra pricing by 20% to $2/$12 per million input/output tokens and GPT-5.6 Luna pricing by 80% to $0.20/$1.20 per million tokens, three weeks after launch. GPT-5.6 Sol pricing is unchanged.
Gemini Robotics 2 extends the original Gemini Robotics system beyond arm-and-hand manipulation to full-body humanoid control, using a multi-model architecture and a new safety benchmark called ASIMOV-Agentic.
GPT Transcribe improves word error rate to 3.31 percent from GPT-4o Transcribe's prior score, and OpenAI cut pricing 25 percent to $0.0045 per minute. The model still trails ElevenLabs Scribe v2 (2.3%), Google Gemini 3 Pro (2.9%), and Mistral Voxtral Small (3%) on the AA-WER benchmark.
Lyria 3.5 improves melodic complexity, lyric quality and prompt adherence, vocal realism and pronunciation, and adds finer control over tempo and track duration, according to Google.
Liquid AI released LFM2.5-Encoder-230M and LFM2.5-Encoder-350M, bidirectional encoder models converted from LFM2.5 decoder backbones, offering 8,192-token context with claimed 3.7x faster CPU inference than ModernBERT-base at long context.
Ai2 launched the OlmoEarth Platform, infrastructure for fine-tuning, evaluating, and running large-scale inference with its OlmoEarth Earth observation foundation models. The platform claims continent-scale processing in about a day using a three-stage CPU/GPU/CPU pipeline and a custom satellite imagery metadata index.
Alibaba released Qwen3.7 Flash, a vision-language reasoning model with a 1M token context window, priced at $0.03 per 1M input tokens and $0.13 per 1M output tokens.
Microsoft's first cybersecurity-specialized model, designed to find vulnerabilities in complex codebases and power the MDASH harness. Launched alongside Perception, an agentic platform for automated security operations.
Composer 2.5 is Cursor's low-cost worker model, built on Kimi K2.5 according to Cursor founder Michael Truell, and priced at $0.50/$2.50 per million input/output tokens. Cursor claims it performs at a level comparable to Opus 4.7 and GPT-5.5 in agentic coding tasks when paired with a frontier planner model.
Claude Opus 5 succeeds Opus 4.8 as Anthropic's flagship model, with the company claiming performance comparable to rival Fable 5 at roughly half the cost. Anthropic says it leads on coding and business-workflow benchmarks while requiring about 85% less intervention than Fable 5.
Google introduced Gemini Omni Flash Preview, a native multimodal model that generates 720p videos with synchronized audio from text, image, and video inputs. This is a new preview release, not an update to an existing model.
Fugu Ultra v1.1 claims performance gains of up to 7.9 points over v1.0, driven largely by improvements on ProgramBench and TerminalBench 2.1, and adds a Claude Code-compatible terminal endpoint. Sakana claims the update outperforms Anthropic's Fable 5 despite Fable 5 not being part of the router's model pool, though all figures are unverified.
Google expanded access to Gemini Spark, an agentic assistant powered by Gemini 3.5, to all US Google AI Pro subscribers and most Google AI Ultra subscribers globally. Free-tier users and Ultra subscribers in the EEA, Switzerland, UK, and Nigeria are excluded from this rollout.
Flux 3 adds native audio generation to videos up to 20 seconds long and introduces a unified Self-Flow architecture trained jointly on images, video, audio, and actions. A companion robotics model, Flux-mimic, is being tested at Audi.
InclusionAI released Ling-3.0-flash, a 124B-parameter MoE model activating ~5.1B parameters per token, offered free via OpenRouter with a 262K context window.
Gemini Nano 4, built on Gemma 4, launches on Samsung's Galaxy Z Fold 8, Fold 8 Ultra, and Flip 8, bringing improved multimodal understanding and expanded language support as the on-device foundation for Google's new Gemini Intelligence feature tier.
Qwen-Image-3.0 expands prompt length to 4,500 tokens and improves text rendering fidelity down to ten pixels, targeting practical layouts like infographics and newspaper pages rather than purely aesthetic images. It is currently invite-only via API and unlikely to be released with open weights, unlike its predecessor.
Laguna S 2.1 launches: Poolside's 118B open-weight MoE (8B active) for agentic coding, 1M context, OpenMDW-1.1 license.
Initial release of LongCat 2.0, a sparse MoE model with 1.6T total parameters and over 1M token context window, optimized for coding and agentic tasks.
FLUX.2 represents a major architecture update with a new 32B parameter model combining Mistral-3 24B VLM with rectified flow transformer, new FLUX.2-VAE trained from scratch, 4MP editing support, and 10-image multi-reference capability.
Initial release of 1.14B parameter multilingual embedding model derived from Ministral-3B through two rounds of structured pruning and distillation.
First multimodal Qwen model above 1 trillion parameters with 2.4T parameters total. Claims improved coding and office work capabilities over Qwen3.7-Max.
Kimi K3 represents a major release from Moonshot AI, claiming frontier-level performance while remaining open source. Independent analyses from Arena.ai and Vals AI suggest competitive performance with flagship frontier models.
Kimi K3 achieves top ranking in Arena's front-end coding capability benchmark, positioning as an open-source alternative to closed U.S. models at 50% lower pricing than GPT-5.6 Sol.
New flagship 8B embedding model ranks #1 on RTEB with 78.5% score. Includes efficient 1B variants with 27% error reduction and Blackwell-optimized NVFP4 deployment option.
Initial release of Inkling, a 975B-parameter MoE model with 41B active parameters trained on 45 trillion tokens. Apache-2.0 licensed for commercial use and fine-tuning.
Cosmos 3 Edge is Nvidia's edge-optimized world model for physical AI applications, following the Cosmos 3 release in May 2026. Designed for robots and vision AI agents requiring real-time environmental perception.
First public model release from Thinking Machines Lab. Open-weight mixture-of-experts architecture designed for enterprise customization rather than frontier performance.
Gemma 4 E2B is a new variant optimized for Pixel 10's Tensor G5 TPU, enabling fully offline multimodal AI capabilities including chat, image recognition, and audio transcription.
Version 2.5 update to the KAT-Coder-Air lightweight coding model line with 256K context window support.
Initial release of Bonsai 27B, claiming to be the first 27-billion parameter model capable of running on-device on iPhone with approximately 4GB memory footprint using 1-bit and ternary quantization.
GPT-5.6 Sol introduces adjustable reasoning effort levels across three model sizes, with approximately five to six settings per size. The Ultra setting uses four subagents to accelerate work at Max-level effort.
Muse Spark 1.1 is an updated version of Meta's foundation model, specifically designed for agentic and coding workloads. Released three months after the original Muse Spark.
Version 2.5 release of KAT-Coder-Pro coding model with 256K context window and updated pricing structure through OpenRouter.
New reasoning-optimized variant of GPT-5.6 Terra with pro mode enabled by default for higher-quality responses on complex tasks.
Initial release of unified audio-text model built on Nemotron-Cascade-2-30B-A3B backbone with extended vocabulary for discrete audio tokens and audio encoder.
Compressed variant of Nemotron-3-Super reduced from 120.7B/12.8B active to 75.3B/9.3B active parameters using Iterative Puzzle framework. Achieves 2× throughput and 8× H100 concurrency while maintaining benchmark accuracy.
GPT-5.6 Terra introduces a balanced mid-tier option in the GPT-5.6 series with 1M context window at $2.50/$15 per million tokens, positioned between flagship Sol and cost-efficient Luna tiers. Targets everyday coding, reasoning, and agentic tasks with February 2026 knowledge cutoff.
GPT-5.6 Luna introduces a cost-optimized model in the GPT-5.6 series with 1M context window at $1/$6 per 1M tokens. Released July 9, 2026 with February 2026 knowledge cutoff.
GPT-5.6 Sol Pro launches as a reasoning-optimized variant of GPT-5.6 Sol, automatically applying extended reasoning mode for higher-quality responses on complex tasks.
GPT-Live-1 introduces full-duplex voice architecture allowing ChatGPT to speak and listen simultaneously, with ability to conduct web searches during conversations. Available in full version for paid users and mini version for free users.
GPT-5.6 Luna Pro is a reasoning-enhanced variant of GPT-5.6 Luna with reasoning.mode set to pro for higher-quality responses on complex tasks. Features 1M context window and February 2026 knowledge cutoff.
Initial release of Robostral Navigate, an 8B embodied navigation model trained entirely in simulation using 400,000 trajectories across 6,000 scenes. Achieves 76.6% R2R-CE validation unseen success using only single RGB camera.
DeepSeek-V4-Flash introduces a 284B parameter MoE architecture with 13B activation, 1M context window support, and three distinct reasoning modes. Unsloth provides optimized GGUF quantizations with Q8 at 162GB.
Laguna XS 2.1 improves on XS.2 with a 5.4% gain on SWE-bench Multilingual (63.1% vs 57.7%) and enhanced terminal-style task performance, while maintaining the 33B total parameter count with 3B activated per token.
GPT-Live-1 replaces OpenAI's turn-based voice model with full-duplex architecture that processes audio input and output simultaneously. The model integrates with GPT-5.5 for reasoning tasks and adds real-time translation.
Initial release of Aion-3.0-Mini, a multi-model collaborative system built on DeepSeek for roleplaying and storytelling applications.
Initial release of 2B parameter Conformer-based ASR model achieving 25.87% average WER on Open Universal Arabic ASR Leaderboard, outperforming models up to 30B parameters.
First release of Muse Image, Meta's first AI image generation model from Meta Superintelligence Labs. Introduces Instagram account prompting, claimed QR code generation, and integration across Meta AI app, Instagram, and WhatsApp.
Initial release of Nex-N2-Mini, the smaller model in the Nex-N2 series with 262K context window and open-source weights.
M2.5 is trained specifically for agent-native execution using reinforcement learning on agentic scaffolds, with emphasis on tool-calling, multi-step task decomposition, and long-horizon coding tasks.
Updated version of the original Leanstral model with improved capabilities for Lean 4 proof assistance and code generation. Built as part of the Mistral Small 4 family with mixture-of-experts architecture.
First release of Nemotron-Labs-TwoTower, a block-wise diffusion model built on the Nemotron-3-Nano-30B backbone. Uses dual-tower architecture to generate blocks of tokens in parallel, claiming 2.42× speedup while retaining 98.7% of baseline quality.
Leanstral 1.5 delivers major performance upgrades over the original Leanstral, saturating miniF2F at 100%, solving 587/672 PutnamBench problems, and achieving state-of-the-art results on FATE-H (87%) and FATE-X (34%) while reducing cost to $4 per problem versus $300+ for competitors.
Redeployed version with enhanced safety classifiers that automatically block cybersecurity tasks and revert to Opus 4.8. Available with restricted usage limits after month-long government-mandated takedown.
Initial release of Portugal's first national AI model, built on EuroLLM-9B foundation with European Portuguese datasets, multimodal capabilities, and expanded context window.
June 2026
Nano Banana 2 Lite replaces the original Nano Banana as Google's fastest and cheapest image generator, generating images in 4 seconds at under $0.04 per thousand images. The model prioritizes speed and cost over quality for high-volume developer pipelines.
Google released Gemini 3.1 Flash Lite Image (Nano Banana 2 Lite), positioned as their fastest and cheapest image generation model optimized for velocity and scale.
Introduces 1M token context, 128K max output, removes sampling parameters (temperature, top_p, top_k), and implements new tokenizer generating ~30% more tokens than Sonnet 4.6 for English text.
New lightweight image generation model optimized for speed and cost, generating 1K images in approximately 4 seconds while maintaining character consistency and editing capabilities of the Nano Banana family.
Claude Sonnet 5 is Anthropic's first Sonnet model of their latest generation, delivering near-Opus intelligence at Sonnet pricing. Key improvements include stronger multi-step reasoning, improved code navigation, and enhanced agentic reliability for production workflows.
Google DeepMind released Nano Banana 2 Lite as the fastest, most cost-efficient image model in the Nano Banana family, designed for high-volume developer workflows. It delivers 4-second text-to-image generation at $0.034 per 1K-resolution image, replacing the original Nano Banana model.
Major release introducing hybrid attention architecture with 90% KV cache reduction, 1M token context window, three reasoning modes, and trained on 32T+ tokens. Includes both Pro (1.6T params) and Flash (284B params) variants.
Initial release of Ornith-1.0, an MIT-licensed agentic coding model built on Gemma 4 and Qwen 3.5, available in 9B, 31B, 35B MoE, and 397B MoE variants.
LFM2.5-230M is a new compact model in the LFM family, built on LFM2 architecture with extended pre-training and reinforcement learning. It features 230M parameters, tool use capabilities through distillation from LFM2.5-350M, and optimized edge inference performance.
First release of Un-0, an image generation model running on simulated oscillator hardware. Produces results comparable to diffusion models like Stable Diffusion but runs on software simulation of non-existent hardware.
DeepSeek-V4-Fable is a distilled variant of Claude-5-Fable built on DeepSeek-V4-Flash, trained specifically for offensive security research using 80,000 CTF trajectories and GRPO reinforcement learning.
Initial release of Qwen-AgentWorld-35B-A3B, a language world model trained through CPT, SFT, and RL stages specifically for agentic environment simulation across seven unified domains.
Initial release of Fugu Ultra, the higher-performance model in Sakana AI's Fugu family using learned multi-agent orchestration rather than a monolithic architecture.
Third behavioral update to GPT-5.5 Instant focused on conversational quality, with improved intent recognition, constraint handling, and recommendation capabilities.
OCR 4 adds bounding boxes, block classification, and inline confidence scores to document extraction. The model expands language coverage to 170 languages and introduces single-container self-hosted deployment.
Initial release of Krea 2 family with Raw (base) and Turbo (post-trained) checkpoints. Turbo variant includes fine-tuning and distillation for faster 8-step generation.
Initial release of Unlimited-OCR, a 3B parameter OCR model building on Deepseek-OCR with support for single-page and multi-page document parsing.
Updated version achieves 85.6% on CyberGym benchmark, surpassing Anthropic's Mythos 5. Released alongside expanded international partnerships.
Initial release of Laguna M.1, a 225B parameter MoE model with 23B active parameters designed for agentic coding tasks, featuring 256 experts and 262K token context window.
Initial release of Mistral OCR API with multimodal document understanding at $1 per 1,000 pages. Claims state-of-the-art accuracy on internal benchmarks across math, tables, and multilingual content.
Initial release of Voxtral TTS, Mistral's first text-to-speech model with 4B parameters supporting 9 languages and voice cloning from minimal audio samples.
First release of Leanstral, a specialized model for Lean 4 proof assistant with 6B active parameters from 120B total. Apache 2.0 licensed with free API endpoint and Mistral Vibe integration.
Devstral 2 powers Mistral Vibe 2.0 with support for custom subagents, multi-choice clarifications, and slash-command skills. API access moved from free to paid pricing.
Mistral OCR 3 achieves 74% win rate over OCR 2 with major improvements in handwriting, forms, scanned documents, and complex tables. Priced at $2 per 1,000 pages ($1 with Batch API).
Initial release of Nano Banana Pro extends the original Nano Banana model with significantly improved multimodal reasoning, real-world grounding, and professional-grade image generation capabilities including 2K/4K outputs.
Gemini 3.1 Flash Image (Nano Banana 2) introduces image generation and editing capabilities with claimed Pro-level quality at Flash model speed and cost efficiency.
First major release of Mistral Large 3, a sparse MoE model with 675B total parameters. Includes multimodal capabilities, multilingual support, and ranks #2 on LMArena among OSS non-reasoning models.
First release of Mistral's code-specialized embedding model with 8192 token context, flexible dimensions, and optimized performance on code retrieval benchmarks including SWE-Bench Lite.
Mistral Medium 3 is a new model introduced alongside Le Chat Enterprise, positioned between Mistral Small and Large tiers. No technical specifications or benchmarks were disclosed.
Pixtral Large is Mistral's new multimodal model, significantly more powerful than Pixtral 12B for document and image understanding. Powers le Chat's document analysis capabilities.
Initial release of Cohere's North family. First agentic coding model from Cohere featuring sparse MoE architecture with 30B total parameters and 3B active.
NVIDIA quantized Google DeepMind's DiffusionGemma 26B A4B IT from 16-bit to 4-bit (NVFP4) using Model Optimizer, reducing memory requirements while maintaining benchmark performance within 1% of the full-precision baseline.
Initial release of FastContext-1.0 repository explorer family with 4B and 30B parameter variants trained via supervised fine-tuning and reinforcement learning.
DiffusionGemma converts the existing Gemma-4-26B-A4B model into a text diffusion model using under 10% of the original training token budget, rather than training from scratch. It trades some quality on standard benchmarks for significantly higher generation speed and new self-correction capabilities.
Gemma 4 family launches on Amazon Bedrock with three variants: 31B dense model, 26B-A4B mixture-of-experts model, and E2B compact model. All variants support multimodal input and built-in reasoning mode.
Kimi K2.7 Code builds on K2.6 with substantial improvements in real-world long-horizon coding tasks and 30% reduction in thinking token usage. Added experimental video input support and improved performance across coding and agentic benchmarks.
Initial release of M3 with 428B parameters, native multimodal training, and MiniMax Sparse Attention enabling 1M context with 15× decode speedup over M2.
Third-generation Apple Foundation Model with 20B parameters using sparse activation. First on-device Apple model to exceed 3B parameters and support native multimodal capabilities.
Google releases DiffusionGemma 26B as open-weight model under Apache 2 license, bringing diffusion-based text generation to production with 500+ tokens/second inference speed.
Initial release of DiffusionGemma, a discrete diffusion-based text generation model built on Gemma 4 26B A4B MoE architecture with encoder-decoder design for parallel token generation.
New speech-to-speech translation model with continuous audio generation across 70+ languages, expanding from Google Meet's previous 5-language support to enable 2000+ language combinations without requiring English as intermediary.
First public release of Anthropic's Mythos model line. Introduces extended-execution code generation with up to 12-hour continuous operation on complex specifications.
Third-generation AFM introduces flash-based inference architecture enabling 20B-parameter on-device model. Uses expert pruning to load subset of parameters into memory while keeping full model in flash storage.
Fable 5 was shut down by government order on June 12, 2026, just 3 days after launch. Expected to return with potential access restrictions after government security review.
Initial release of North Mini Code 1.0, a 30B-parameter sparse MoE model with 3B active parameters, trained specifically for agentic coding tasks with tool-use capabilities.
Initial release of Nex-N2-Pro, a 397B parameter MoE model with 17B active parameters and 262K context window, available free via OpenRouter.
Apple's most advanced cloud-based AI model, running on Nvidia GPUs in Google cloud infrastructure with privacy guarantees. Refined using Google Gemini frontier model outputs.
Initial release of Nemotron 3 Ultra, a 550B-parameter MoE model with 55B active parameters, hybrid Transformer-Mamba architecture, and 1M context window designed for agentic workflows.
First open-weight release from Ideogram featuring 9.3B parameters, trained from scratch with structured JSON prompting, native 2K support, and state-of-the-art text rendering capabilities.
First release of Nemotron-3-Ultra, NVIDIA's 550B parameter (55B active) frontier model with hybrid LatentMoE architecture combining Mamba-2, MoE, and Attention layers. Features 1M token context and toggle-able reasoning capabilities.
Nemotron 3.5 ASR expands the English-only Nemotron 3 ASR to support 40 language-locales from a single 600M-parameter checkpoint with cache-aware streaming architecture.
Initial release of Nemotron-3-Ultra, a 550B parameter model trained December 2025-April 2026 with hybrid LatentMoE architecture, 1M token context, and configurable reasoning capabilities.
Nemotron 3.5 adds custom policy enforcement, unified multimodal evaluation, auditable reasoning traces, and releases the training dataset. Built on Gemma 3 4B IT with 128K context window.
Gemma 4 introduces encoder-free architecture in the 12B Unified model, processing all modalities directly through a single decoder-only transformer. The family spans five models from 2.3B to 30.7B parameters with extended context windows up to 256K tokens.
Qwen3.7 Plus introduces multimodal capabilities to the Qwen3.7 series with text and image input support, featuring a 1 million token context window at competitive pricing.
First mid-sized Gemma model with native audio support, eliminates multimodal encoders for direct vision and audio processing through LLM backbone. Requires only 16GB RAM for local inference.
Gemma 4 12B is a new 12-billion-parameter multimodal model designed for local inference on consumer laptops. Google claims it matches the performance of their 26-billion-parameter mixture-of-experts model while running on devices with 16GB RAM.
First open-weight release of Ideogram's text-to-image model with 9.3B parameters, structured JSON prompting, and native 2K resolution support.
Anthropic expands restricted preview to 150 additional organizations across 15+ countries, focusing on critical infrastructure sectors including utilities and healthcare. Public release remains delayed pending safety measures.
Initial release of Microsoft's first advanced reasoning model, trained from scratch on proprietary data without distillation from other models.
Initial release of Cosmos3-Nano, a 16B-parameter omnimodal world model for Physical AI applications with 256K context window and support for generating video, audio, images, and robot actions from multimodal inputs.
Initial release limited to Project Glasswing partners. Public Mythos-class models planned for release within weeks of June 2, 2026.
Holo3.1 adds mobile automation support, native function-calling, and quantized checkpoints for local inference. AndroidWorld performance improves from 67% to 79.3%, with 2× end-to-end speedup on DGX Spark.
Initial release of Mellum2-12B-A2.5B-Thinking, a reasoning-augmented model trained with supervised fine-tuning and reinforcement learning with verifiable rewards on math-heavy data.
Initial release of MAI-Thinking-1, a 35B parameter reasoning model trained on commercially licensed data without third-party model distillation. Available to select early partners.
M3 introduces MiniMax Sparse Attention to enable 1M-token context at approximately 1/20th the compute cost of previous generation. Native multimodal training on interleaved data with interactive user-simulator tuning.
Initial GGUF release of Step-3.7-Flash with seven quantization variants from BF16 (394GB) to IQ3_XXS (76GB), all optimized for local deployment on consumer hardware with 128GB memory.
First release of Cosmos 3, a unified omni-model that combines video generation, physical reasoning, and action prediction in a single Mixture-of-Transformers architecture, replacing the previous fragmented Cosmos model suite.
May 2026
Initial release of Cosmos3-Super-Text2Image as part of NVIDIA's Cosmos3 omnimodal world model collection, featuring 64B parameters and Mixture-of-Transformers architecture for Physical AI applications.
Initial release of Cosmos 3, NVIDIA's omnimodal world foundation model platform for Physical AI, featuring 64B-parameter variants with Mixture-of-Transformers architecture supporting video, image, audio, and robot action generation.
Opus 4.8 shows substantially lower rates of misaligned behavior compared to 4.7, with approximately 4x improvement in catching code flaws. Fast mode API pricing reduced by 3x.
Initial release of Step-3.7-Flash, a 198B-parameter sparse MoE vision-language model with 256K context window, three reasoning levels, and production-focused architecture delivering 400 tokens/sec.
Mistral Large 3 is Mistral's first MoE model since Mixtral, featuring 675B total parameters with 41B active. Released with both base and instruction-tuned versions under Apache 2.0, with multimodal image understanding and optimized inference support.
Mistral OCR 3 introduces major improvements over OCR 2 with claimed 74% win rate on forms, handwriting, scanned documents, and complex tables. Pricing set at $2 per 1,000 pages ($1 with batch API).
First reasoning model release from Mistral AI with 24B parameters and native multilingual chain-of-thought capabilities. Open-sourced under Apache 2.0 license.
Initial release of Mistral Saba, a 24B-parameter model specialized for Arabic and South Asian languages including Tamil, designed for regional deployment.
Major release unifying reasoning, multimodal, and coding capabilities. First Mistral model with configurable reasoning effort parameter and native multimodal support in the Small family.
Codestral 25.08 improves code completion with 30% higher acceptance rates, 10% better retention, and 50% fewer runaway generations. Chat mode gains 5% on instruction-following and code benchmarks.
SDK adds support for Claude Opus 4-8 model with mid-conversation system blocks and granular output token usage details.
Initial release of Voxtral family with 24B (Small) and 3B (Mini) variants under Apache 2.0. Features 32K context, native multilingual support, and direct function-calling from speech.
New release of Devstral Medium achieving 61.6% on SWE-Bench Verified. Companion release of Devstral Small 1.1 (24B parameters) scores 53.6% and is open-sourced under Apache 2.0.
Fast-mode variant of Claude Opus 4.8 offering higher output speed at 2x pricing compared to the standard version.
Music v2 adds mid-track genre switching, section-based composition, and targeted editing capabilities. Released 10 months after Music v1.
Initial release of Gemini Omni Flash, the first tier of Google's multimodal video generation model with avatar cloning and physics modeling capabilities.
Initial release of LocateAnything-3B with Parallel Box Decoding architecture, trained on 12M images across natural scenes, robotics, driving, GUI, and document domains.
Initial release of Stable Audio 3 Medium with 2B parameters, supporting variable-length audio generation up to 6+ minutes with sub-2-second inference times on H200 GPU.
Initial release of diffusion language model family trained on 1.3T pretraining tokens and 45B fine-tuning tokens. Supports autoregressive, diffusion, and self-speculation generation modes with up to 6.4× speedup over traditional AR models.
DeepSeek permanently reduced V4 Pro pricing by 75%, dropping input tokens to $0.003625 per million and output tokens to $0.87 per million. Previously promotional pricing is now permanent.
Gemini 3.5 Flash now powers Google's AI Mode search, adding support for multimodal inputs including images, video files, and Chrome tabs with improved intent anticipation.
Mistral Medium 3.5 merges instruction-following, reasoning, and coding into a 128B dense model with 256k context. Released as open weights under modified MIT license with configurable reasoning effort and new vision encoder.
Hy-MT2 represents a major version release with new 1.8B, 7B, and 30B-A3B model sizes supporting 33 languages, extreme quantization via AngelSlim, and the IFMTBench benchmark for instruction-following evaluation.
Initial release of Command A+ open source model featuring 25B active parameters in a 218B parameter sparse mixture-of-experts architecture with vision, tool use, and reasoning capabilities.
Flagship release of Qwen3.7 series with 1M token context window and agent-first design. Notable improvements in coding and agentic performance over prior Qwen generations with explicit prompt caching support.
Microsoft released Fara1.5-27B, a 27B-parameter vision-only computer use agent fine-tuned from Qwen3.5-27B, supporting 262K context and MIT licensing. The model introduces built-in 'critical points' safety pausing for irreversible actions and is designed for deployment inside Microsoft's MagenticLite sandbox.
Initial release of Grok Build 0.1, xAI's first coding-specialized model with 256K context window designed for agentic workflows and CLI integration.
Gemini 3.5 Flash is Google's first model in the 3.5 series, claiming improved agentic and coding capabilities at reduced cost compared to frontier models. Features enhanced safety measures with reasoning checks before responses.
OlmoEarth v1.1 reduces compute costs by up to 3x through token sequence length reduction, collapsing Sentinel-2 resolution bands into single tokens while maintaining similar benchmark performance to v1.
Initial release of Qianfan-OCR-Fast, a specialized OCR model with 66K context window and domain-specific training for improved document processing performance.
Complete architecture rebuild from XLM-RoBERTa to ModernBERT, expanding context from 512 to 32K tokens (64x increase) and adding code retrieval. The 97M model achieves 60.3 on MTEB Multilingual Retrieval (+12.2 over R1), highest in its size class.
First release of Gemini Omni family supporting multimodal video generation with conversational editing. Speech and audio editing capabilities withheld pending safety testing.
Fast-mode variant of Claude Opus 4.7 with identical capabilities but prioritized output speed at 6x premium pricing ($30/$150 per 1M tokens vs standard rates).
Initial release of Perceptron Mk1, a vision-language model specializing in video understanding and spatial annotation with optional reasoning capabilities.
Initial release of Trinity Large Thinking as a free, open source reasoning model with 262K context window and focus on agentic workloads.
Gemma 4 E4B assistant introduces Multi-Token Prediction architecture for speculative decoding, achieving up to 2x inference speedup. Features 4.5B effective parameters with Per-Layer Embeddings optimized for on-device deployment.
Initial release of Ring-2.6-1T, a 1T parameter thinking model with 63B active parameters, featuring adaptive reasoning and 262K context window optimized for agent workflows.
Google launches Flash Lite variant of Gemini 3.1 at 50% cost reduction with maintained 1M context window and four-level reasoning system.
First OpenAI voice model with GPT-5-class reasoning capabilities, designed for live voice interactions with tool calling and natural conversation flow.
GPT-5.5-Cyber is a variant of GPT-5.5 with relaxed safeguards specifically for vetted cybersecurity teams. The model is trained to be more permissive on security-related tasks compared to the standard GPT-5.5 release.
Major release introducing Gemma 4 family with four model sizes (E2B, E4B, 26B A4B MoE, 31B dense), Multi-Token Prediction drafters for 2x speedup, extended context windows up to 256K, enhanced multimodal capabilities, and improved reasoning performance.
Initial release of Multi-Token Prediction assistant model for Gemma 4 26B A4B, enabling up to 2x inference speedup through speculative decoding while maintaining identical output quality.
Initial release of CoBuddy code generation model with 131K context window, native tool calling, and reasoning support, available for free on OpenRouter.
Granite Speech 4.1 2B introduces dual-head CTC encoder, frame importance sampling, improved multilingual ASR accuracy, and punctuation/truecasing across all languages. Two new variants add speaker attribution with timestamps and non-autoregressive architecture.
GPT-5.5 Instant replaced GPT-5.3 Instant as the default ChatGPT model with significant accuracy improvements and more concise response formatting. OpenAI claims 52.5% fewer hallucinations on high-stakes prompts and reduced use of emojis and unnecessary formatting.
IBM released Granite 4.1 family in 3B, 8B, and 30B sizes under Apache 2.0 license. Unsloth released 21 GGUF quantized variants of the 3B model.
April 2026
Granite 4.1 8B is a new release in IBM's Granite 4.1 family, offering an 8B-parameter dense decoder-only model with enterprise-focused capabilities. Released under Apache 2.0 license with 131K context window and multilingual support.
R2 upgrades architecture from XLM-RoBERTa to ModernBERT, extends context from 512 to 32,768 tokens, expands vocabulary to 262K tokens, and adds Matryoshka dimension reduction support. Performance improves by 11.8 points on MTEB Retrieval.
Granite 4.1 30B introduces enhanced tool-calling, improved instruction following through updated SFT-RL pipeline, and 131K context window. Released under Apache 2.0 license with competitive performance on code and reasoning benchmarks.
Second-generation XS size model with enhanced tool calling and reasoning capabilities, quantized to fp8 for production efficiency.
Initial release of 13-billion-parameter vintage language model trained exclusively on public domain texts published before 1931.
First omni-modal release in Nemotron 3 line, adding audio and video capabilities to previous vision-language model. Uses new 30B-A3B MoE architecture with hybrid Mamba-Transformer design.
Initial release of Nemotron 3 Nano Omni, a 30B-parameter multimodal MoE model designed as a perception sub-agent for enterprise systems. Features hybrid Transformer-Mamba architecture with specialized video processing and extended reasoning capabilities.
Laguna XS.2 introduces a 33B parameter MoE architecture with 3B activated parameters per token, featuring mixed sliding window and global attention layers, native reasoning support, and optimization for local deployment.
Initial release of Nemotron 3 Nano Omni, a 31B-parameter MoE multimodal model with video, audio, image, and text understanding, 256K context window, and dedicated reasoning mode with chain-of-thought capabilities.
Initial release of Owl Alpha, OpenRouter's first foundation model designed specifically for agentic workloads with native tool use and 1M+ context window.
Initial release of Nemotron-3-Nano-Omni-30B-A3B, a multimodal MoE model with 31B parameters combining video, audio, image, and text understanding with reasoning capabilities.
Initial release of Nemotron 3 Nano Omni, a multimodal MoE model with 30B total parameters (3B active) combining video, audio, image, and text understanding in a single inference pass with 131K token context.
Moonshot AI introduced a router endpoint that automatically redirects to the most current model in the Kimi family, featuring a 262,144 token context window.
Initial release of GPT Mini Latest with 400,000 token context window and auto-redirect to newest GPT Mini family version.
Google released a dynamic router endpoint that automatically redirects to the latest Gemini Pro model, featuring 1,048,576-token context and reasoning capabilities.
Initial release of Qwen3.6 Max Preview, a proprietary 1 trillion parameter sparse MoE model with 262K context window and integrated thinking mode for agentic workflows.
Qwen3.6 Flash introduces multimodal support for text, images, and video with a 1M token context window. Features tiered pricing structure with base rates of $0.25/$1.50 per 1M tokens for prompts under 256K tokens.
MiMo-V2.5-Pro introduces 1.02T total parameters (up from 310B in V2.5) with 42B active parameters, extends context to 1M tokens, and adds hybrid attention architecture with sliding window and global attention patterns. The model achieves 99.6% on GSM8K and maintains coherence at extreme context lengths.
New sparse MoE model with 35B total parameters but only 3B active per token, featuring 262K context window, multimodal support, and integrated reasoning mode.
Updated Qwen3.5 Plus with 1M token context window and tiered pricing above 256K tokens.
Google released Gemini Flash Latest as a dynamic router that automatically redirects to the newest Gemini Flash model, featuring 1,048,576 token context and reasoning capabilities.
Qwen3.6 27B introduces video processing capabilities alongside existing text and image support, with a 262K context window and built-in thinking mode for agentic coding and reasoning tasks.
Major release with 1.6T total parameters, 1M token context window, and substantial efficiency improvements over V3.2. DeepSeek claims near-frontier performance at a fraction of the cost.
DeepSeek-V4-Flash is a new 284B-parameter MoE model with 13B activated parameters, featuring hybrid attention architecture that reduces inference costs by 73% at million-token context lengths. Introduces three reasoning effort modes and achieves competitive performance with frontier models on coding and mathematical reasoning.
GPT-5.5 Pro introduces a 1M+ token context window optimized for complex reasoning tasks, agentic coding, and multi-step workflows with multimodal support.
Major version release with claimed competitive performance against leading US models. DeepSeek emphasizes improved coding capabilities and domestic chip compatibility.
Initial release of Ling-2.6-1T, a 1 trillion parameter instruct model with 262K context window and fast-thinking architecture designed for cost-efficient agent deployments.
Qwen3.6-27B adds extended context to 1M+ tokens, introduces thinking preservation feature, and shows 2-5% improvements in coding agent benchmarks over Qwen3.5-27B. FP8 quantization enables efficient deployment with claimed near-identical performance to full precision.
Initial release of Trinity Large Preview, a 400B-parameter sparse MoE model with 13B active parameters per token, featuring up to 512K context window support and optimization for agentic workflows.
Gemma 4 E2B demonstrated running as a vision-language agent on NVIDIA Jetson Orin Nano Super (8GB), autonomously deciding when to access webcam based on conversational context with no hardcoded triggers.
Initial release of Privacy Filter, a 1.5B-parameter bidirectional token classifier for detecting 8 PII categories with 128K context window and Apache 2.0 license.
Qwen3.6-27B is a 27B dense model that claims flagship-level coding performance surpassing the 397B Qwen3.5-397B-A17B. Available at 55.6GB full size or 16.8GB quantized.
Initial preview release of Hy3, Tencent's Mixture-of-Experts model with configurable reasoning modes designed for production agentic workflows.
Major update to OpenAI's image generation model with significantly improved quality, maximum resolution of 3840x2160 pixels, and pricing at $30 per million output tokens.
Initial release of Pareto Code Router with dynamic model selection based on min_coding_score parameter (0-1 scale) and 200K context window.
GPT-5.4 Image 2 combines GPT-5.4 reasoning capabilities with image generation from GPT Image 2, supporting text, image, and file inputs with a 272K token context window.
Major update to OpenAI's image generation capabilities with significantly improved text rendering and UI element generation accuracy.
Initial release of Ling-2.6-flash with 104B total parameters (7.4B active), featuring 262K context window and optimized for agent applications with fast response times.
First release of Kimi K2.6 introduces 1T-parameter MoE architecture with agent swarm capabilities, 256K context, and competitive performance on SWE-Bench and agentic benchmarks.
Initial release of Qianfan-OCR-Fast, a specialized OCR model with 65K context window and free pricing. Claims performance improvements over Qianfan-OCR.
GR00T N1.7 upgrades to Cosmos-Reason2-2B VLM backbone and adds EgoScale pre-training on 20,854 hours of human egocentric video, improving dexterity and generalization over N1.6.
First open-weight release of Qwen3.6 series, featuring improved agentic coding capabilities, thinking preservation for context retention, and FP8 quantization. Built on community feedback to prioritize stability and real-world utility.
Python SDK v0.96.0 adds Claude Opus 4-7 support, introduces token budgets for cost management, and user profiles for personalized interactions.
Opus 4.7 introduces significantly more aggressive safety guardrails that automatically detect and block requests related to cybersecurity uses, resulting in a 10-15x increase in reported false positive refusals.
Initial preview release of Gemini 3.1 Flash TTS, Google's first prompt-controlled text-to-speech model that accepts theatrical-style direction for voice characteristics, accents, and delivery style.
Major release introducing multi-modal 3D world generation capabilities, replacing video-based outputs with persistent 3D assets. Centered on WorldMirror 2.0, a 1.2B parameter model for reconstruction.
Gemini 3.1 Flash TTS introduces audio tags for granular control over vocal style, pace, and delivery through natural language commands. The model achieves an Elo score of 1,211 and includes mandatory SynthID watermarking.
Fine-tuned variant of GPT-5.4 specifically built for defensive cybersecurity work with reduced restrictions on security-related tasks. Initial access limited to verified security professionals through Trusted Access for Cyber program.
Anthropic released Mythos, a security-focused AI model designed to identify and exploit zero-day vulnerabilities. The model remains unreleased publicly due to security concerns.
Initial release of Trinity-Large-Thinking, a 400B parameter open-weight reasoning model with 256 mixture-of-experts (13B active per token) optimized for agent tasks. Apache 2.0 licensed, trained on 17 trillion tokens over 33 days on 2,048 Nvidia B300 GPUs.
Refreshed version of LFM2-VL-450M with updated LFM2.5-350M backbone. Adds bounding box prediction and function calling capabilities while improving performance across vision and language benchmarks.
Released HY-Embodied-0.5 suite with MoT-2B (2.2B active parameters, 4B total) and 32B variants. MoT-2B trained on 200B+ tokens of embodied data outperforms similarly-sized competitors on 22 embodied benchmarks.
Amazon Bedrock now enables supervised fine-tuning, reinforcement fine-tuning, and model distillation for Nova 2 Lite. Fine-tuned models deploy on-demand at standard inference pricing without provisioned capacity.
NVIDIA released Alpamayo 2 Super, a new 34B-parameter vision-language-action model combining a 32B VLM backbone with a 2.3B diffusion-based action decoder for autonomous vehicle perception, planning, and reasoning tasks.
Mythos is Anthropic's new frontier model, positioned as larger and more intelligent than its Opus models. It is deployed exclusively through Project Glasswing for defensive cybersecurity work with 40+ vetted partner organizations, with claims of identifying thousands of zero-day vulnerabilities during early testing.
Claude Mythos Preview demonstrates unprecedented capability in autonomous vulnerability discovery and exploitation, finding decades-old bugs in OpenBSD, FFmpeg, and FreeBSD. Deployment restricted to 11-partner coalition for defensive cybersecurity use only.
Amazon Nova 2 Sonic enables real-time conversational podcast generation with 1M token context window and native support for seven languages through Amazon Bedrock.
Claude Mythos Preview released under restricted access through Project Glasswing. Model demonstrates exceptional cybersecurity research capabilities including discovery of 27-year-old OpenBSD TCP SACK vulnerability and Linux privilege escalation flaws.
Initial release of Harrier embedding model. Trained on 2B+ examples with GPT-5 synthetic data. Achieves top ranking on MTEB v2 multilingual benchmark with 131K context window.
Claude Mythos is Anthropic's specialized model for cybersecurity vulnerability discovery, designed to identify critical flaws in operating systems, browsers, and software. The model shows improvements over Claude Opus 4.6 in reasoning, agent-based capabilities, and coding.
Gemma 4 E4B adds multimodal capabilities (text, image, audio), extended 128K context window, native reasoning modes, and function-calling support compared to Gemma 3. Achieves 69.4% MMLU Pro with 4.5B effective parameters optimized for mobile and edge deployment.
Initial release of Bonsai 8B 1-bit quantized model. Achieves 14x compression with claimed competitive performance on standard benchmarks. Also released Bonsai 4B and Bonsai 1.7B variants.
Gemma 4 introduces multimodal support (text, image, video, audio on small models), extended context windows (128K-256K tokens), configurable reasoning modes, and native function calling. Available in four sizes with both dense and MoE architectures.
Tencent released OmniWeaving, an open-source unified video generation model with reasoning capabilities and compositional video creation. Built on HunyuanVideo-1.5, it supports eight video generation tasks and introduces IntelligentVBench benchmark.
Zhipu AI released GLM-5V-Turbo, adding multimodal capabilities to its GLM-5 series. The model generates code from design mockups and video inputs while maintaining text-only coding performance, integrating directly with Claude Code and OpenClaw agents.
Gemma 4 26B A4B uses Mixture-of-Experts with 3.8B active parameters for efficient inference. Features 256K context window, multimodal input (text/image), native reasoning modes, and function-calling for agentic workflows.
Qwen3.6 Plus introduces hybrid linear attention with sparse mixture-of-experts routing, achieving 78.8 on SWE-bench Verified. Major improvements in coding, reasoning, and multimodal capabilities over 3.5 series.
NVIDIA released NVFP4-quantized version of Google DeepMind's Gemma 4 31B IT model optimized for consumer GPU inference. Maintains 256K context window and multimodal capabilities with <0.5% performance degradation on reasoning and coding benchmarks.
Google DeepMind introduces Gemma 4 31B with multimodal input (text and images), 256K context window, configurable reasoning mode, and native function calling. Free release under Apache 2.0 license.
Gemma 4 introduces multimodal capabilities (text, image, video, audio on small models), extended 256K context windows, configurable reasoning modes, and hybrid dense/mixture-of-experts architectures. Substantial improvements in coding benchmarks, long-context reasoning, and on-device deployment efficiency compared to Gemma 3.
Gemma 4 introduces multimodal support, 256K context window, Apache 2.0 permissive licensing, and mixture of experts variant. First major version update with explicit focus on enterprise deployment without data usage restrictions.
Gemma 4 introduces multimodal capabilities with native image and audio support, extended 128K context window, built-in reasoning modes with configurable thinking, and hybrid attention architecture combining local and global attention for efficiency.
Gemma 4 introduces multimodal capabilities (text, image, video support), reasoning modes, 256K context windows, and Mixture-of-Experts architecture. The 26B A4B variant uses sparse activation for near-dense-31B performance with 4B-model inference speed.
Microsoft released MAI-Transcribe-1, a speech-to-text model achieving lowest FLEURS benchmark word error rate at 2.5x faster inference than Azure Fast. Priced at $0.36 per audio hour, supporting 25 languages and challenging recording conditions.
Gemma 4 introduces four model sizes (2B-31B) with improved reasoning and agentic capabilities. Apache 2.0 licensing replaces previous restrictions. 31B model ranks #3 on Arena AI leaderboard.
Qwen 3.6 Plus introduces a hybrid architecture with linear attention and sparse mixture-of-experts routing, delivering major improvements in agentic coding, front-end development, and reasoning over the 3.5 series. Achieves 78.8 on SWE-bench Verified.
Initial release of Falcon Perception 0.6B early-fusion Transformer for open-vocabulary grounding and segmentation. Introduces Chain-of-Perception output interface and PBench diagnostic benchmark with five capability levels.
Meta's first closed-source AI model with planned paid developer access, marking a strategic shift from the open-source Llama series. Released from Meta Superintelligence Labs under new AI leadership.
Holo3-122B-A10B released with 78.85% OSWorld score using mixture-of-experts architecture (122B total, 10B active parameters). Trained via agentic learning flywheel with synthetic data augmentation and curated reinforcement learning. Holo3-35B-A3B variant open-sourced under Apache 2.0.
March 2026
Microsoft released the Harrier-OSS embedding model family with three parameter sizes (270M, 600M, 27B) supporting multilingual inputs, 32K token context, and knowledge distillation techniques. The 27B variant achieves 74.3 on MTEB v2 benchmark.
Google released Veo 3.1 Lite, a cost-optimized video generation model priced at less than 50% of Veo 3.1 Fast. Designed for high-volume applications with same generation speed as Veo 3.1 Fast.
Granite 4.0 3B Vision introduces a compact vision-language model optimized for enterprise document processing with DeepStack Injection architecture and ChartNet dataset training. Shipped as LoRA adapter on Granite 4.0 Micro for modular text-only fallback support.
Qwen3.5-Omni expands from Qwen3-Omni with 8x context window increase (32K to 256K tokens), 6x language support expansion (11 to 74 languages), hybrid attention-MoE architecture, and ARIA token interleaving for improved real-time speech synthesis. Demonstrates emergent code-generation capability from spoken and video input.
Grok 4.20 Multi-Agent is a specialized variant designed for collaborative agent-based workflows with parallel agent coordination. Scales agent count based on reasoning effort: 4 agents at low/medium effort, 16 agents at high/xhigh effort.
Google releases Lyria 3 Pro Preview, a music generation model producing full-length songs with vocals, lyrics, and instrumental arrangements. Priced at $0.08 per song through the Gemini API with 1M token context window.
Qwen 3.6 Plus Preview introduces a hybrid architecture with 1M token context window, improved reasoning and agentic behavior over 3.5 series. Available free on OpenRouter with data collection for model improvement.
Microsoft releases Harrier-OSS-v1 family of multilingual embedding models in three sizes (270M, 0.6B, 27B parameters) trained with contrastive learning and knowledge distillation. The 0.6B variant achieves 69.0 MTEB v2 score with 32,768 token context window and supports 45+ languages.
Lyria 3 Clip Preview introduces Google's music generation model to the Gemini API with clip-based pricing at $0.04 per 30-second generation.
KAT-Coder-Pro V2 builds on earlier KAT-Coder versions with enhanced agentic coding capabilities for large-scale production environments and added web design generation for landing pages and presentation decks.
Cohere releases Transcribe, a 2B parameter open-source speech recognition model with 5.42% WER on the Hugging Face leaderboard, supporting 14 languages under Apache 2.0 license.
Dreamina Seedance 2.0 launches in CapCut with IP safeguards including face-detection blocks, unauthorized IP generation prevention, and invisible watermarking to identify AI-generated content.
NVIDIA released gpt-oss-puzzle-88B, an inference-optimized 88B-parameter mixture-of-experts model derived from gpt-oss-120B using the Puzzle NAS framework. Achieves 1.63× throughput improvement on long-context and up to 2.82× on single H100s while maintaining parent accuracy through heterogeneous expert pruning, selective window attention, and knowledge distillation with RL optimization.
Google released Gemini 3.1 Flash Live, an audio-focused model optimized for multilingual conversations. The model powers the expansion of Search Live to 200+ countries with claimed improvements to response speed and conversation naturalness.
Gemini 3.1 Flash Live improves upon 2.5 Flash Native Audio with enhanced acoustic recognition, background noise filtering, lower latency, and extended conversation context.
Expanded from 30-second to 3-minute song generation with improved structural composition control. Added support for specifying discrete song elements.
Apple developed RubiCap, a rubric-guided reinforcement learning framework for dense image captioning that achieves state-of-the-art results with 2B-7B parameter models, outperforming competitors up to 72B parameters.
Initial release of MolmoWeb with 4B and 8B parameter variants. Includes full training dataset (MolmoWebMix), model weights, and evaluation tools under Apache 2.0 license.
Nemotron 3 Super now available on Amazon Bedrock as fully managed serverless inference. 120B parameter MoE model with 12B active parameters, 256K context, claims 5x throughput improvement and 2x accuracy gain over previous version.
Composer 2 launched with frontier-level coding intelligence, built on Moonshot AI's open-source Kimi 2.5 model with additional reinforcement learning training applied by Cursor (75% of final compute).
Xiaomi released MiMo-V2-Pro as 3x larger successor to MiMo-V2-Flash (Dec 2025), reaching 1T parameters with 42B active per request. Benchmarks place it 3rd globally on PinchBench/ClawEval, nearly matching Claude Opus 4.6 on coding (78% vs 80.8%) while costing 80% less per input token.
M2.7 introduces autonomous participation in its own development through 100+ self-optimization rounds, achieving 30% performance improvement on internal coding tasks and competitive benchmark scores against leading Western models.
MAI-Image-2 improves upon MAI-Image-1 with enhanced photorealism, natural lighting, and notably adds reliable text rendering capabilities for practical design applications. The model ranks third on Arena.ai leaderboard, up from ninth place for the previous version.
Composer 2 achieves 61.3 on CursorBench (+38% vs Composer 1.5) through improved pretraining and reinforcement learning on long-horizon tasks. Pricing set at $0.50/$2.50 per 1M tokens, undercutting Claude and GPT-4 by 60-90%.
MiMo-V2-Pro is Xiaomi's flagship foundation model launch featuring 1T+ parameters and 1M context window optimized for agent systems and complex workflow orchestration.
Xiaomi debuts MiMo-V2-Omni, a frontier omni-modal model processing image, video, and audio natively. 262K context with strong agentic capabilities including visual grounding and code execution.
Xiaomi launches MiMo-V2-Pro, their flagship 1T-parameter foundation model with 1M context. Optimized for agentic scenarios, ranking among global top tier on standard benchmarks.
Qwen3.5-Max-Preview launches as Alibaba's largest model (1T+ params) in the Qwen3.5 series. 262K context with thinking mode (82K CoT). Beats previous flagship on reasoning, multilingual, and agentic tasks. $1.20/$6.00 per 1M tokens.
GPT-5.4 mini introduces major improvements in coding, reasoning, and computer control capabilities over GPT-5 mini. Model runs over 2x faster and achieves near-full-GPT-5.4 performance on multiple benchmarks while consuming 30% of quota in agentic systems.
GPT-5.4 mini, OpenAI's fastest variant of GPT-5.4, is now generally available in GitHub Copilot. The model claims to be the highest-performing mini offering for coding tasks.
GPT-5.4 Nano launches as the smallest, fastest member of the GPT-5.4 family. 400K context, multimodal input, optimized for high-volume agentic tasks at $0.20/$1.25 per 1M tokens.
Mistral Small 4 launches unifying Magistral reasoning, Pixtral multimodal, and agentic coding into one model. 262K context at $0.15/$0.60 per 1M tokens.
Initial release of Nemotron-3-Nano-4B-GGUF, a quantized (Q4_K_M) 4B parameter edge model with hybrid Mamba-2 architecture. Supports controllable reasoning modes and 262K context window for edge AI applications including gaming NPCs and local voice assistants.
Initial release of MiniMax M2.7 — next-gen LLM with multi-agent collaboration, 204K context, SWE-Pro 56.2%, Terminal Bench 2 57.0%
GLM 5 Turbo launches with 203K context and fast inference optimized for agent-driven environments. Improved complex reasoning over base GLM 5 at $0.96/$3.20 per 1M tokens.
ByteDance releases Seedream 4.5, their latest image generation model with major quality improvements over Seedream 4.0.
Nvidia released Llama 3.1 Nemotron 70B Instruct, an instruction-tuned variant of Meta's Llama 3.1 70B model optimized for developer applications.
Minimax released M1 80k, expanding its M1 model family with an 80,000-token context window for extended document processing.
Google launched Ask Maps, integrating Gemini AI into Google Maps to allow users to ask complex contextual navigation questions. The chatbot personalizes responses based on user search history and saved locations.
Mistral AI releases Pixtral Large, a new multimodal model supporting image and text inputs with 128K context window.
Minimax releases M1 40k model with 40,000-token context window. Initial release with limited publicly disclosed specifications.
Nvidia releases Nemotron 3 Super, a 120B hybrid MoE model with 1M context window, latent expert routing, and multi-token prediction. Fully open-weight under NVIDIA Open License.
ByteDance launches Seed-2.0-Lite, a cost-efficient multimodal enterprise model with 262K context. Strong agent capabilities at $0.25/$2.00 per 1M tokens.
NVIDIA releases Nemotron-3-Super-120B, a 120B parameter model with latent MoE architecture optimized for conversational tasks across 8 languages.
NVIDIA releases Nemotron-3-Super-120B-A12B-BF16, a 120 billion parameter model with latent MoE architecture for efficient text generation across 8 languages.
StepFun releases Step-3.5-Flash-Base as an open-source text generation model optimized for efficient inference under Apache 2.0 license.
Qwen3.5-9B released as multimodal 9-billion parameter model supporting image and text inputs. Available under Apache 2.0 license on Hugging Face.
Gemini 3.1 Flash Lite Preview launches as Google's high-efficiency model for high-volume use. 1M context at $0.25/$1.50, outperforms Gemini 2.5 Flash Lite.
Initial release of Qwen3.5-2B, a 2-billion-parameter multimodal model supporting image and text processing.
Qwen3.5-0.8B released as an 800-million-parameter multimodal model for edge inference. Supports image and text inputs under Apache 2.0 licensing.
Qwen3.5-4B released as a 4 billion parameter multimodal model supporting image and text inputs. Apache 2.0 licensed for open-source use.
Initial release of Context-1, a 20B parameter Mixture of Experts retrieval agent model trained for multi-hop search with self-editing context capabilities.
Released FP8-quantized version of Qwen3.5-35B-A3B, reducing memory requirements while maintaining multimodal capabilities. Compatible with Transformers endpoints and Azure deployment.
Gemini 3.1 Flash-Lite achieves 2.5x faster first-token latency than Gemini 2.5 Flash with 360 tokens/second throughput. Output pricing increased to $1.50 per million tokens from $0.40.
February 2026
Nano Banana 2 (Gemini 3.1 Flash Image Preview) debuts as Google's fastest image generation model. Pro-level quality at Flash speed with $0.50/$3.00 per 1M tokens.
Seed-2.0-Mini launches targeting latency-sensitive scenarios with 262K context and four reasoning effort levels at $0.10/$0.40 per 1M tokens.
Qwen3.5-35B-A3B-Base released as a 35-billion parameter multimodal model with Apache 2.0 license. Part of the Qwen3.5 mixture-of-experts family.
Qwen3.5-Flash debuts with 1M context and ultra-low $0.065/$0.26 pricing. Hybrid architecture delivers a leap in inference efficiency over Qwen 3 series.
Mercury 2 launches. The first reasoning-capable diffusion LLM. Hits 1,009 tokens/sec on NVIDIA Blackwell with 1.7s end-to-end latency — 5x faster than speed-optimized autoregressive models. Tunable reasoning, native tool use, schema-aligned JSON output. OpenAI-compatible API.
Qwen3.5-35B-A3B released as open-weight multimodal model with 35B parameters. Apache 2.0 licensed, supports image and text inputs with conversational capabilities.
Qwen3.5-27B released as a 27-billion parameter multimodal model supporting image-text-to-text tasks. Available under Apache 2.0 license with transformer endpoint compatibility.
Cohere Labs releases tiny-aya-global, a multilingual text generation model fine-tuned from tiny-aya-base to support conversational tasks across 100+ languages including major and low-resource languages.
Gemini 3.1 Pro enters public preview in GitHub Copilot with focus on efficient edit-then-test loops and agentic coding capabilities.
Lyria 3 integrated into Gemini, enabling 30-second music track generation with vocals, lyrics, and cover art from text prompts or uploaded media.
Alibaba released Qwen 3.5 series claiming performance parity with proprietary frontier models while optimized for commodity hardware, directly challenging closed-source AI model economics.
Claude Opus 4.6 — major GPQA and reasoning improvements; ARC-AGI-2 jump from 37.6% to 68.8%.
GLM-5.1 introduces sustained agentic reasoning over hundreds of iterations with improved performance on SWE-Bench Pro (58.4%, +3.3pp vs. GLM-5) and NL2Repo (42.7%, +6.8pp vs. GLM-5). The model maintains productivity across longer problem-solving sessions through iterative experimentation and strategy revision.
Qwen3.5 397B A17B launches as a hybrid linear-attention + sparse MoE vision-language model. 262K context at $0.39/$2.34 per 1M tokens with state-of-the-art performance.
Qwen3.5 Plus launches with 1M context and strong cross-domain performance in academia, finance, marketing, programming, and science at $0.30/$1.20 per 1M tokens.
Gemini 3.1 Pro released as upgrade to Gemini 3 Pro. Enhanced reasoning for complex multi-step problems. Preview access via AI Studio.
MiniMax-M2.5 launches.
Claude Sonnet 4.6 brings improved multi-step reasoning and stronger agentic task performance over 4.5, at the same price.
GPT-5.3-Codex launches. First model combining the Codex and GPT-5 training stacks — best-in-class code generation plus general-purpose reasoning in one model, ~25% faster than its predecessor.
Step-3.5-Flash launches.
Grok 4 mini brings next-generation reasoning at low cost. Significantly outperforms Grok 3 mini on AIME and competitive coding tasks.
January 2026
Kimi K2.5 launches as Moonshot AI's native multimodal model with state-of-the-art visual coding and agent swarm paradigm. 262K context at $0.45/$2.20 per 1M tokens.
Claude Sonnet 4.5 January patch with stability improvements for long agentic sessions and better tool use across multi-turn workflows.
Gemini 2.5 Pro preview January update with longer stable thinking output windows and improved factual grounding on complex queries.
Global rollout suspended following copyright cease-and-desist letters from Disney and Paramount Skydance over use of copyrighted training material.
December 2025
DeepSeek V3 December update with improved instruction following and expanded Chinese-English code-switching performance.
Gemini 3 Flash launches. Fast, cost-efficient Gemini 3 workhorse with near-Pro reasoning at a fraction of the cost. 1M context. Rolled out globally in the Gemini app, Search AI Mode, and the API on day one.
Gemini 3.0 Pro debuts as first Gemini 3 generation model. Upgraded reasoning and native multimodal understanding. Quickly superseded by 3.1.
Mistral Small 3.1 December update with improved vision accuracy and expanded support for structured data extraction from images.
November 2025
FLUX.2 launches. Second-generation FLUX with 4MP photoreal output, multi-reference character consistency, accurate text rendering, and editing in one checkpoint. Open-weights dev variant alongside pro/flex API tiers.
xAI releases Grok 4.1 Thinking variant with reasoning-focused capabilities. Technical specifications and pricing not yet disclosed.
Kimi K2 Thinking debuts as Moonshot AI's most advanced open reasoning model. Trillion-parameter MoE with 32B active params for agentic long-horizon reasoning.
Gemini 2.5 Flash November preview with configurable thinking budget improvements and better cost efficiency at higher thinking token counts.
Claude 3.7 Sonnet November update with significantly improved agentic task completion rates and more reliable computer use across complex workflows.
October 2025
Added audio embedding support to Nova Multimodal Embeddings. Model now processes audio content alongside text, images, documents, and video with four embedding dimension options (3,072, 1,024, 384, 256) using Matryoshka Representation Learning.
MiniMax-M2 launches.
Veo 3.1 launches. Veo upgrade with richer native audio, better prompt adherence and image-to-video, plus editing controls in Flow: ingredients-to-video, frame interpolation, and scene extension.
Claude Opus 4.5 — improved agentic coding and reasoning.
Baidu launches ERNIE 4.5 21B A3B Thinking, an upgraded lightweight MoE reasoning model. Top-tier on math, science, and coding benchmarks at $0.07/$0.28 per 1M tokens.
Llama 4 Scout patch fixing multimodal tokenization issues and improving throughput for long-context document tasks.
September 2025
DeepSeek-V3.2-Exp launches. Experimental release introducing DeepSeek Sparse Attention (DSA), cutting long-context inference cost dramatically — API prices dropped 50%+ overnight. The bridge architecture to V4.
Qwen3-Max launches. Alibaba's first 1T+ parameter model, released at Apsara Conference 2025. Top-tier coding and agentic performance (69.6 SWE-bench Verified); closed-weights API flagship of the Qwen3 line.
Grok 4 Fast launches. Cost-efficient Grok 4 variant with a 2M-token context and unified reasoning/non-reasoning weights — Grok-4-level benchmark scores using ~40% fewer thinking tokens.
Mistral Large 2 September update with improved code generation and expanded function calling support for complex tool schemas.
Claude 3.7 Sonnet September patch with improved computer use stability and reduced hallucination rate in extended thinking mode.
Qwen3-Next-80B-A3B launches. Architecture preview for Qwen3.5: hybrid attention (Gated DeltaNet + Gated Attention) with ultra-sparse MoE — 80B params, only 3B active. 10x cheaper to train and faster at long context than Qwen3-32B.
Kimi K2 September update brings improvements to the trillion-parameter MoE model with 256K long-context inference at $0.40/$1.60 per 1M tokens.
August 2025
Nano Banana (Gemini 2.5 Flash Image) launches. The viral image editing model that topped LMArena as 'nano-banana' before Google revealed it as Gemini 2.5 Flash Image. Best-in-class character consistency and multi-image fusion. Superseded by Nano Banana Pro and 2.