LLM News

Every LLM release, update, and milestone.

0
model release

Qwen Releases Qwen3.8 2.4T A95B, a 2.4-Trillion-Parameter Open-Weight MoE Model

Qwen has released Qwen3.8 2.4T A95B, an open-weight sparse mixture-of-experts model with 2.4 trillion total parameters and 95 billion active parameters per forward pass. The model is the open-weight variant of Qwen3.8 Max, targeting coding, research, complex reasoning, and agentic workflows with a 262K token context window.

2 min readvia openrouter.ai
0
product updateGoogle DeepMind

Google DeepMind Launches SL2T, a Sign-Language-to-Text Model Trained on 100,000+ Hours Across 50+ Languages

Google DeepMind has released SL2T, a massively multilingual sign-language-to-text translation model trained on over 100,000 hours of data across 50+ sign languages. The model powers new sign-to-text dictation features in Gboard and Live Transcribe on Pixel 11, starting with American Sign Language to English.

0
product update

Mistral Adds EU/US Regional Routing and Paid Priority Queue, Both With Coverage Gaps

Mistral has made regional inference generally available, letting customers route requests through EU or US servers for a 10 percent surcharge, while also launching a paid Priority Tier that charges 1.75x standard pricing for faster processing during peak traffic. Both offerings carry significant limitations on what data and features they actually cover.

0
changelogAnthropic

Anthropic Adds Machine-Readable Watermarks to Claude-Generated Text and Files

Anthropic is adding machine-readable watermarks to Claude-generated text and digital signatures to generated files to comply with the EU AI Act's Article 50 transparency mandate. The change applies to models launched after Aug. 2 and rolls out globally, though Anthropic admits detection can fail on heavily edited or short text.

2 min readvia axios.com
0
researchAnthropic

Researchers Extract Hidden Chain-of-Thought from OpenAI, Anthropic, Google Models via Shared Encryption Keys

A paper published at stolen-thoughts.com demonstrates that encrypted reasoning traces returned by OpenAI, Anthropic, and Google APIs used the same encryption key across models in a family, allowing attackers to jailbreak weaker sibling models into revealing a stronger model's hidden chain-of-thought in plaintext. All three providers have since patched the vulnerability.

0
model releaseNVIDIA

NVIDIA Releases Nemotron 3.5 Lightning 30B-A3B: 3B-Active MoE Model With 1M-Token Context, Quantized for Single-GPU Depl

NVIDIA has published NVIDIA-Nemotron-3.5-Lightning-30B-A3B-NVFP4, a 30-billion-parameter Mixture-of-Experts model with only 3B active parameters, a hybrid Mamba-2/MoE/Attention architecture, and support for up to 1 million tokens of context. The NVFP4-quantized checkpoint is designed to run on a single DGX Spark (GB10) or H100 GPU.

2 min readvia huggingface.co