LLM News

Every LLM release, update, and milestone.

0
research

Tavus says 48% of testers mistook its Griffin video AI for a real person on a one-minute call

Tavus has introduced Griffin, which it calls the first 'Human Interaction Model' for real-time face-to-face video conversation. In a Tavus study, 48% of participants believed Griffin was a real person after a one-minute call, versus a 2% maximum for earlier systems. A limited research preview, Griffin-Lite, is open only to select testers.

0
researchAnthropic

Graphite: Opus 5.5 uses 'this matters' 116x more than humans as AI writing tells persist

Marketing firm Graphite identified 13,000 phrases that appear at least twice as often in AI-generated writing as in human writing. Claude Opus 5.5 uses "this matters" 116 times more than humans, while OpenAI's Astra favors "corrective framing" more than 100 times as often. Em-dash use has collapsed across frontier models, but total tells are holding steady, according to Graphite.

4 min readvia techcrunch.com ↗
0
model release

Ideogram 4.5 launches with native 2K output and four tiers from $0.008 to $0.22 per image

Ideogram has released Ideogram 4.5, an image model it claims edits only the area a user specifies and leaves the rest untouched. It offers four quality tiers from 0.8 to 22 cents per image, all at native 2K resolution, via the Ideogram platform and API. An open-weight release is promised but not yet dated.

0
model releaseAmazon Web Services

Amazon open-sources Strands Decider 2B, a small decision model built on a Qwen3.5-2B base

Amazon Web Services has released Strands Decider 2B, an open-source model that chooses among pre-decided options and returns a confidence score instead of generating text. It is inspired by TypeSafe's Jev and is small enough to run locally. Amazon says it briefly topped the Jevbench ranking for models of its size.

3 min readvia techcrunch.com ↗
0
product updateAmazon Web Services

AWS details ambient agent pattern on Bedrock AgentCore: S3 events trigger jobs, one ask_human tool pauses for approval

AWS published a reference implementation for ambient agents on Amazon Bedrock AgentCore. S3 uploads or scheduled events create jobs that an agent runs, pausing for human input through a single ask_human tool. Each agent turn is capped at the 15-minute Lambda timeout.

0
product updateAmazon Web Services

uniopen lifts Amazon Nova 2 Lite moderation F1 from 0.585 to 0.855 using LoRA fine-tuning on SageMaker AI

Taiwan retail platform uniopen adapted Amazon Nova 2 Lite to its two-axis moderation policy using LoRA supervised fine-tuning in Amazon SageMaker AI, plus a prompt-format change. According to AWS, Per Behavior Macro F1 rose from 0.5852 to 0.8550 and Subject Type Macro F1 from 0.4162 to 0.8491, both above production targets.

0
researchAi2

Ai2 releases Olmo-core 3, an open MoE training stack benchmarked at 1.2T parameters

Ai2 released Olmo-core 3, an open training framework for large mixture-of-experts models. The company reports about 2.7× the throughput of its earlier FSDP-based implementation and benchmarks up to 1.2 trillion total parameters on 512 NVIDIA B300 GPUs. It is the infrastructure for Ai2's next MoE-based Olmo, not a model release.

3 min readvia huggingface.co ↗
0
model releaseUnbiased

Unbiased releases Pareto 26.10 Preview: 1M context, $0.80/$3.20 per 1M tokens on OpenRouter

Unbiased has listed Pareto 26.10 Preview on OpenRouter, a multimodal composite model with a 1.0M-token context window priced at $0.80 input and $3.20 output per 1M tokens. The company says it targets research, coding, and agentic workflows, and warns the preview may change without notice. No benchmark scores have been published.

3 min readvia openrouter.ai ↗
0
product updateOpenAI

OpenAI says it shut down a 15,000-account campaign to extract model reasoning; the attack still worked on Azure

OpenAI says it detected and shut down an adversarial distillation campaign involving more than 15,000 accounts, linked in part to people associated with Moonshot AI. Researchers report that the same reasoning-extraction attack still worked on Microsoft Azure on September 13 against every OpenAI model they tested, including GPT-6 Astra.

4 min readvia the-decoder.com ↗
0
model release

Google launches Gemini 4 Argon at $2/$10 per 1M tokens, rolling out first to Fairwind cybersecurity partners

Google has launched Gemini 4 Argon, the first model in the Gemini 4 family, at an introductory $2 per million input tokens and $10 per million output tokens. Artificial Analysis says it matches GPT-6 Astra on its Intelligence Index at 60% of the cost per task, with a 15% hallucination rate. Access starts with Google's Fairwind Program for governments and trusted partners.

4 min readvia engadget.com ↗
0
model release

Google DeepMind Releases Gemini 4 Argon, Expands Output Limit to 1M Tokens

Google DeepMind has released Gemini 4 Argon, a frontier model built for long-horizon reasoning with an industry-leading 1 million output token limit. The model is rolling out first to trusted cyber defenders through Google's Fairwind Program, with pricing set at $2 per million input tokens and $10 per million output tokens.

3 min readvia deepmind.google ↗