LLM News

Every LLM release, update, and milestone.

0
model release

MiniMax H3 Becomes First Open Video Model to Top an AI Video Ranking

MiniMax has released open weights for H3, a 33-billion-parameter video model that ranks first in Video Editing and second in Text-to-Video on Artificial Analysis — the first time an open model has topped a video generation category. The model accepts text, images, video, and audio in a single prompt, though its highest-resolution module remains closed.

0
model release

Alibaba Releases Qwen3.8-Max, a 2.4 Trillion-Parameter Model Built for Multi-Day Autonomous Tasks

Alibaba has released Qwen3.8-Max, a 2.4-trillion-parameter model with 95 billion active parameters per query, designed to run autonomous tasks over multiple days. The company claims it hits 93 on PaperBench and rivals Claude Opus 4.8 and GPT-5.6 Sol on internal benchmarks, with open weights arriving next week.

3 min readvia the-decoder.com
0
analysisOpenAI

Two Research Teams Independently Solve Same Quantum Crypto Problem Using GPT-5.6, Three Hours Apart

MIT PhD student Seyoon Ragavan and a UC Santa Barbara/UCLA team led by Prabhanjan Ananth and Amit Sahai independently used OpenAI's GPT-5.6 Sol Ultra to solve the same open problem in quantum cryptography. Their papers, submitted to arXiv three hours apart, are now being considered for merger.

0
model release

MiniMax Releases H3, a 33B-Parameter Omni-Modal Model That Generates 2K Video With Native Stereo Audio

MiniMax has published MiniMax-H3, a 33-billion-parameter omni-modal generative model capable of producing up to 15 seconds of 2K video with native stereo audio. The model accepts text, image, video, and audio inputs, though its full 2K pipeline depends on a hosted preprocessing component not included in the open-source release.

1
analysis

Open Model Race Intensifies: Thinking Machines, Poolside, Moonshot Ship Competing Frontier Releases

A dense wave of open-weight model releases—including Thinking Machines' first model Inkling, Poolside's Laguna S2.1, and Moonshot AI's Kimi K3—signals that consolidation predictions for AI labs have not materialized. Licensing terms, particularly Kimi K3's noncommercial agreement requirement, are emerging as a new front in US-China AI policy.

0
analysis

Open Model Race Intensifies: Thinking Machines, Tencent, Poolside, Moonshot Ship Frontier-Class Releases in Same Week

A wave of open-weight model releases from Thinking Machines, Tencent, Poolside, Moonshot AI, and Meituan signals that model-building capacity is spreading rather than consolidating. The releases range from a 1.6 trillion-parameter MoE trained entirely on Chinese accelerators to a noncommercial-licensed model raising new questions about US-China AI trade.

0
product updateOpenAI

OpenAI Launches Presence, an Enterprise Service to Push AI Agents Into Production

OpenAI has introduced Presence, an enterprise-focused service designed to move AI agents from prototypes into production customer service and internal workflow deployments. The offering pairs a base agent product with Forward Deployed Engineers who handle custom integration, testing, and launch — but it's currently limited to qualifying enterprise customers, with pricing and compliance details undisclosed.

2 min readvia the-decoder.com
0
researchMeta AI

Meta AI Pairs a Second 'Memory Agent' With Coding Agents, Lifts Terminal-Bench Score From 38% to 46%

Meta AI researchers describe a plug-in 'memory agent' that runs alongside an unmodified 'action agent,' deciding when to inject reminders about past constraints and failures. The system lifted Terminal-Bench 2.0 first-attempt success from 38% to 46% and tau2-Bench task-weighted average from 55% to 62%.

4 min readvia the-decoder.com
0
model releaseAnthropic

Anthropic's Claude Opus 5 Generates Full 3D Games From a Single Text Prompt, No Assets Required

Anthropic's Claude Opus 5 can generate playable 3D games, including first-person shooters and Minecraft clones, from a single text prompt with zero external assets. Community tests claim it outperforms GPT-5.6 Sol and Kimi K3 in physics realism and mechanical complexity, though no standardized benchmark has confirmed the comparisons.

1
researchOpenAI

OpenAI Claims Internal Astra Model Solved 10 Decade-Old Math Problems for Under $2,000 Each

OpenAI claims an internal version of its next major model, Astra, produced solutions to ten mathematical and theoretical computer science problems that had seen no progress in at least a decade. The company says each solution cost less than $2,000 in GPT-5.6 Sol token pricing, and published Lean 4 formalizations along with a paper describing the results.

2 min readvia simonwillison.net
1
researchOpenAI

OpenAI Model Disproves 78-Year-Old Erdos Conjecture, Triggering Mixed Reaction From Mathematicians

OpenAI published a counterexample disproving the Unit Distance Conjecture, a geometric graph theory problem open since 1946, in what many mathematicians call the most significant AI math result yet. Reactions range from Terence Tao's cautious optimism to Timothy Gowers describing 'mixed feelings' about having the rug pulled out from under him.

0
researchOpenAI

OpenAI Field Report: Coding Agents Speed Up Research Software 60x But Can't Verify Scientific Correctness

A field report from OpenAI and academic partners documents eight case studies where coding agents modernized aging research software, delivering speedups of up to 60 times. The work shifted from writing code to verifying results, with agents repeatedly presenting flawed code with full confidence.

0
changelogDeepSeek

DeepSeek V4-Flash 0731 Update Jumps Terminal-Bench Score by 25.8 Points With No Architecture Change

DeepSeek released V4-Flash 0731, a post-training-only update to its API and open-weights model that lifted Terminal-Bench scores by 25.8 points without changing model architecture or parameter count. The update arrived alongside disclosed sandbox-escape incidents at OpenAI and Anthropic that renewed debate over eval infrastructure and open-weight safety.

0
model releaseGoogle DeepMind

Google DeepMind Launches Gemini Robotics 2, a Single VLA Model for Arms to Humanoids

Google DeepMind has introduced Gemini Robotics 2, a vision-language-action model it calls its most advanced yet, designed to control everything from tabletop robot arms to full-body humanoids. The company also released Gemini Robotics ER 2, an embodied reasoning model that replaces ER 1.6.