LLM News

Every LLM release, update, and milestone.

0
product updateAnthropic

Cline CLI v3.0.69 raises MCP startup timeout from 3 to 10 seconds, fixing silently dropped Windows servers

Cline CLI v3.0.69 raises the default MCP server startup timeout from 3 seconds to 10 seconds. On Windows, servers launched via npx or uvx were being silently dropped. The release also fixes Claude requests through custom Anthropic base URLs, reasoning-level mismatches, and a broken `cline config --json` command.

3 min readvia github.com ↗
0
product updateAnthropic

Anthropic launches Claude for Google Workspace add-on in public beta, adding sidebars to Docs, Sheets and Slides

Anthropic has released the Claude for Google Workspace add-on in public beta, placing a Claude sidebar inside Google Docs, Sheets, and Slides. It is available to all paid Claude users through the Google Workspace Marketplace, and includes an "ask before edits" preview mode.

2 min readvia 9to5google.com ↗
0
model releaseGoogle DeepMind

Google DeepMind releases EmbeddingGemma 2: 740M-parameter open embedding model spanning text, image, video, audio

Google DeepMind has released EmbeddingGemma 2, an Apache 2.0 open embedding model with 740M total parameters that maps text, images, video and audio into a single 768-dimensional vector space. According to the model card, it improves code retrieval on MTEB (code, v1) from 68.76 to 78.68 over its predecessor while keeping an 8,192-token context window.

0
model release

Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0

Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says it runs in ~191MB of active RAM for text-only weights and ~567MB for the full multimodal model on a quantized Pixel 11 Pro. Google also released a Mac app, AI Edge Foresight, to demonstrate it.

3 min readvia 9to5google.com ↗
0
model release

Google releases EmbeddingGemma 2: 740M-parameter multimodal embedding model under Apache 2.0

Google announced EmbeddingGemma 2, a 740M-parameter natively multimodal embedding model built on the Gemma 4 architecture and released under Apache 2.0. Google says the quantized model needs about 191MB of active RAM for text-only weights and about 567MB for the full multimodal model on a Pixel 11 Pro. Google also launched a Mac app, AI Edge Foresight, to demonstrate it.

3 min readvia 9to5google.com ↗
1
model release

Google releases Nano Banana 2.1 image model: $1.50/$30 per 1M tokens, 66K context

Google's Nano Banana 2.1 (Gemini Nano Banana 2.1) is an image generation and editing model on the Flash tier, listed on OpenRouter at $1.50 input and $30 output per 1M tokens with a 66K context window. It supports 1K, 2K, and 4K output and succeeds Nano Banana 2 and Nano Banana Pro, according to the listing.

0
product updateAmazon Web Services

AWS publishes reference build for voice airline concierge using Nova 2.5 Sonic and Bedrock AgentCore

AWS has published a reference architecture for a voice travel concierge that pairs Amazon Nova 2.5 Sonic with Bedrock AgentCore runtime, AgentCore Gateway over MCP, and Bedrock Knowledge Bases. It deploys with a single AWS CDK script and runs against a sample airline backend with synthetic data. The post does not disclose pricing, context window, or benchmark figures for Nova 2.5 Sonic.

1
model releaseMistral AI

Mistral releases Large 4, a 1-trillion-parameter multimodal model, with open weights due in three weeks

Mistral AI released Mistral Large 4 (ML4), a multimodal model with one trillion parameters, on Tuesday. It is currently available only through a public guardrail endpoint, and Mistral plans to publish the weights in about three weeks after safety testing. Benchmark results, pricing and context window have not been disclosed.

0
model release

Mistral Large 4 enters public preview: 1T-parameter open-weight multimodal model, weights due by end of October

Mistral AI has launched a public preview of Mistral Large 4, a 1-trillion-parameter natively multimodal model with 49 billion active parameters. The preview API is live on Mistral Studio, and open weights are promised by the end of October 2026. Pricing and context window have not been disclosed.

3 min readvia mistral.ai ↗
0
model release

Reflection unveils 501B-parameter Beam, Mistral previews 1T-parameter Large 4, both open-weight

Reflection introduced Beam, a 501B-parameter mixture-of-experts model with 23B active parameters. Mistral said it is finishing Mistral Large 4, a 1T-parameter multimodal model with 49B active parameters. Both companies plan open-weight releases in October, and both are positioning the models against Chinese open-weight leaders.

3 min readvia axios.com ↗
0
model release

Reflection announces Beam, a 501B-parameter open-weight coding model with 23B active parameters

Reflection has announced Beam, its first open-weight model, a 501B-parameter mixture-of-experts system with 23B active parameters per token, built for coding, reasoning and agentic tasks. The company claims it matches GLM 5.2 on demanding reasoning tasks with three to four times less compute. Weights are due under Apache 2.0 later this month.

4 min readvia the-decoder.com ↗
0
changelogOpenAI

OpenAI ships opt-in textGrain text watermarking in API, with EU ChatGPT and Codex rollout to follow

OpenAI has launched textGrain, an invisible statistical text watermark, as opt-in for API customers worldwide on select models. ChatGPT and Codex output in the EU will be watermarked in the coming weeks, in response to the EU AI Act. OpenAI says detection drops from about 92% to 17% when 25% of words in a 400-token passage are replaced.

3 min readvia 9to5mac.com ↗
Page 1 of 51Next →