LLM News

Every LLM release, update, and milestone.

0
changelogOpenAI

OpenAI ships opt-in textGrain text watermarking in API, with EU ChatGPT and Codex rollout to follow

OpenAI has launched textGrain, an invisible statistical text watermark, as opt-in for API customers worldwide on select models. ChatGPT and Codex output in the EU will be watermarked in the coming weeks, in response to the EU AI Act. OpenAI says detection drops from about 92% to 17% when 25% of words in a 400-token passage are replaced.

3 min readvia 9to5mac.com ↗
0
product updateOpenAI

OpenAI to watermark ChatGPT and Codex text in the EU under AI Act; API opt-in available worldwide

OpenAI will add an invisible watermark to text generated by ChatGPT and Codex in the European Union to comply with the EU AI Act's transparency rules. Developers anywhere can enable it on select API models starting today, but it is off by default. OpenAI's own tests show detection falling from about 92% to 66% after 10% of words are replaced with synonyms.

3 min readvia techcrunch.com ↗
0
model release

Reflection AI unveils Beam: 501B-parameter open-weight MoE with 1M-token context

Reflection AI has unveiled Beam, a text-only mixture-of-experts model with 501 billion total parameters, 23 billion active, and a 1 million token context window. The company claims it matches Z.ai's GLM-5.2 on advanced reasoning benchmarks while using 3-4x less inference compute. Weights and the full technical report are due later this month.

3 min readvia techcrunch.com ↗
0
product updateAnthropic

Claude Opus 5.5 and Sonnet 5.5 now on Amazon Bedrock in AWS GovCloud (US), with Claude Code support

Claude Opus 5.5 and Claude Sonnet 5.5 are available on Amazon Bedrock in AWS GovCloud (US) Regions. AWS published a setup guide for running Anthropic's Claude Code against them for regulated workloads, including ITAR. Pricing, context window and benchmark figures were not disclosed.

0
product updateAmazon Web Services

AWS ships aws-ai-ml skill so Kiro, Claude Code and Codex can benchmark SageMaker inference endpoints

Amazon SageMaker AI has released the aws-ai-ml skill, distributed through the Agent Toolkit for AWS. It lets MCP-compatible coding agents such as Kiro, Claude Code and Codex benchmark endpoints, recommend deployment configurations and generate executable SageMaker Python SDK v3 code.

3 min readvia aws.amazon.com ↗
0
benchmarkGitHub

GitHub launches ReviewBench, an open benchmark for AI code review agents built on real pull requests

GitHub has launched ReviewBench, an open benchmark for evaluating AI code review agents. According to GitHub, it uses representative pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. Dataset size, scores, and licensing were not included in the announcement excerpt.

2 min readvia github.blog ↗
0
benchmarkGitHub

GitHub launches ReviewBench, an open benchmark for AI code review agents built on real pull requests

GitHub has launched ReviewBench, an open benchmark for evaluating AI code review agents. According to GitHub, it uses representative pull requests, multi-source ground truth, calibrated evaluation, and production-aligned metrics. Dataset size, model scores, and licensing details were not included in the announcement summary.

3 min readvia github.blog ↗
0
product updateAmazon Web Services

AWS releases open-source MCP server to automate cross-account promotion of Amazon Quick agents

AWS has published the Quick Resource Migrator, a sample MCP server on Amazon Bedrock AgentCore that promotes Amazon Quick resources between AWS accounts in a single tool call. It is idempotent, never deletes from the target, and writes versioned S3 backups before every update.

0
research

Google's RRSI cuts overfitting in self-improving agents: up to 4.7-point gains on unseen tasks with ~30% fewer tokens

Google Cloud AI Research and several universities introduced RRSI, a method that stops self-optimizing agent harnesses from memorizing their test tasks. According to the paper, it gains up to 4.7 points on five unseen benchmarks and uses about 30% fewer runtime tokens than the unregularized version, with the underlying model frozen.

4 min readvia the-decoder.com ↗
0
benchmarkAleph Alpha

Aleph Alpha benchmark: Chinese AI models balanced on just 17-41% of 967 sensitive-topic prompts

An Aleph Alpha study of 967 politically sensitive prompts found that only 17 to 41 percent of responses from Alibaba's Qwen, DeepSeek and Moonshot's Kimi were balanced, according to the company's own AI scorer. DeepSeek V4 Pro refused roughly two-thirds of questions. Aleph Alpha sells "sovereign AI" to governments, which gives it a commercial interest in the result.

3 min readvia the-decoder.com ↗
0
model release

China Telecom's Xing4.0-29B-A4B: 29B MoE, 4B Active, 256K Context, Trained Fully on Ascend NPUs

China Telecom AI's Xing4.0-29B-A4B (formerly the TeleChat line) is a mixture-of-experts model with 29B total and 4B active parameters and a native 256K context window, extensible to 512K. The company claims it is the first model of this scale trained entirely on Ascend NPUs with MindSpore. Community GGUF quantizations from Venastine-Research are already available.

3 min readvia huggingface.co ↗
0
researchGoogle DeepMind

DeepMind essay argues AGI will emerge from human-agent networks, not a lone superintelligence

Google-affiliated researchers Benjamin Bratton, Blaise Agüera y Arcas and James Manyika propose "Artificial Symbiotic Intelligence," a framework in which AGI emerges from a social system of people and AI agents rather than a single self-improving machine. The essay, written for the Deepmind Institute, is a conceptual argument and reports no benchmark results.

0
product updateAnthropic

Anthropic adds Mods to Claude Code, a plugin system that hooks into tool calls, prompts and UI rendering

Anthropic released Mods for Claude Code, a plugin system built on JavaScript and TypeScript functions that hook into events such as tool calls, user prompts and UI rendering. Mods are not sandboxed and run with the user's permissions. They work in the CLI, the desktop app and, partly, the VS Code extension.

3 min readvia the-decoder.com ↗