LLM News

Every LLM release, update, and milestone.

0
model release

Z.ai Releases GLM-5.3-Prime, a High-Throughput Variant of GLM-5.3 with 1M-Token Context

Z.ai has released GLM-5.3-Prime, a high-speed variant of its GLM-5.3 model that delivers 1.5-2x the output throughput through inference acceleration while retaining the full 1M-token context window. The model is priced at $2.80 per 1M input tokens and $8.80 per 1M output tokens, targeting coding and long-horizon agentic workloads.

2 min readvia openrouter.ai ↗
0
product update

Amazon Blocks Meta's Muse AI Agent From Making Purchases, Citing Terms of Service

Amazon has blocked Meta's new Muse AI personal agent from completing purchases on its platform, saying the tool violates its terms of service. The move comes as Muse tops Apple's App Store and Meta stock rallies more than 20% in two weeks, with Wall Street watching Zuckerberg's Meta Connect keynote for signs of a broader agentic AI platform strategy.

3 min readvia cnbc.com ↗
0
product updateAmazon Web Services

AWS Adds Open Weight Models to Amazon Bedrock for Terminal-Based Coding Agents via OpenCode

Amazon Bedrock now supports the open source coding agent OpenCode paired with open weight models including Moonshot AI's Kimi K3, OpenAI's GPT-OSS 120B, and NVIDIA's Nemotron 3 Super 120B. The setup keeps inference inside a customer's AWS account with per-token pricing instead of per-seat subscriptions.

0
research

Google Brings Persistent, Encrypted Memory to Cloud AI Without Breaking On-Device Privacy Guarantees

Google is adding a persistent memory layer to its Private AI Compute platform, letting AI assistants retain context across devices while keeping data encrypted with keys held only on user devices. The company published a technical whitepaper and an independent security audit alongside the update.

3 min readvia deepmind.google ↗
0
model release

Google Launches Gemini 3.8 Flash TTS and Flash-Lite TTS with Voice Creation from Text Prompts

Google DeepMind has released Gemini 3.8 Flash TTS and Gemini 3.8 Flash-Lite TTS, text-to-speech models that generate custom voices from natural language prompts and support line-by-line performance direction. The models top Hume AI's Voice Design Benchmark at 71.4 and claim first and second place on its Overall Quality Index.

3 min readvia deepmind.google ↗
0
model releaseStealth

Anonymous Stealth Model "Space Bunny Alpha" Debuts on OpenRouter With 1M-Token Context, Free During Preview

A previously unknown AI provider has released Space Bunny Alpha, a stealth model on OpenRouter offering a 1M-token context window, adjustable reasoning effort, and multimodal input support. The model is free during its preview period, though its developer remains unnamed.

0
model releaseNVIDIA

NVIDIA Releases Nemotron 3 Diarization, a 100M-Parameter Open-Weight Model Ranked #1 on Voice Arena's Diarization-Bench

NVIDIA released Nemotron 3 Diarization, a 100-million-parameter open-weight model that identifies who is speaking and when in audio conversations. It ranked #1 among 17 system configurations on Voice Arena's Diarization-Bench with a 14.72% diarization error rate, supporting up to eight speakers in both live and recorded audio.

3 min readvia huggingface.co ↗
0
model releaseUpstage

Upstage Releases Solar Mini 4: 35B MoE Model with 524K Context at $0.05/$0.20 per Million Tokens

Upstage has released Solar Mini 4, a compact mixture-of-experts model with 35B total parameters, 3B active parameters, and a 524K token context window. The model targets agentic workloads and is priced at $0.05 per 1M input tokens and $0.20 per 1M output tokens, a promotional 50% discount off standard rates.

2 min readvia openrouter.ai ↗