model release

Mira Murati's Thinking Machines announces full-duplex AI model with 0.40-second response time

TL;DR

Thinking Machines Lab, founded by former OpenAI CTO Mira Murati, announced TML-Interaction-Small, a full-duplex AI model that processes input while generating responses simultaneously. The company claims 0.40-second response time, matching natural human conversation speed.

2 min read
0

Mira Murati's Thinking Machines announces full-duplex AI model with 0.40-second response time

Thinking Machines Lab, the AI startup founded by former OpenAI CTO Mira Murati, announced what it calls "interaction models" — AI systems that process input and generate responses simultaneously rather than in traditional turn-taking fashion.

The company's first model, TML-Interaction-Small, claims a 0.40-second response time, which Thinking Machines says matches natural human conversation speed and is "significantly faster than comparable models from OpenAI and Google." The technical architecture enables full-duplex communication, meaning the model can listen while it speaks, similar to a phone call rather than a text message exchange.

Release timeline and availability

This is a research preview, not a public product. Thinking Machines plans a "limited research preview" within the next few months, with wider release scheduled for later in 2026. No pricing information has been disclosed.

The company has not released benchmark scores, parameter count, context window specifications, or other technical details beyond the response latency claim.

Technical approach

Current AI models follow a strict turn-taking protocol: user input is processed completely before response generation begins. Thinking Machines' approach processes incoming audio or text while simultaneously generating output, which the company argues should be "native to a model, not bolted on."

The distinction matters for applications requiring real-time interaction, such as voice assistants or conversational interfaces where interruption and natural flow are expected.

What this means

Full-duplex communication represents a meaningful architectural shift if the claimed performance holds under real-world conditions. A 0.40-second response time would indeed approach human conversation norms (typical human response latency ranges from 0.2 to 0.6 seconds). However, without independent verification, public testing, or detailed benchmarks, it's impossible to assess whether this translates to genuinely better user experience or represents a marginal improvement over existing streaming architectures. The value proposition depends entirely on whether simultaneous processing creates noticeably more natural interactions than current streaming implementations, which remains unproven until researchers and users can test the system directly.

Related Articles

model release

NVIDIA Releases Nemotron VoiceChat 11B, an Open Full-Duplex Speech Model with Live Tool Calling

NVIDIA has released NemotronLabs VoiceChat 11B, an 11-billion-parameter end-to-end full-duplex speech model that unifies streaming speech understanding and generation in one architecture. The model claims to be the first open full-duplex system to support live tool calling during natural conversation, with ~450ms turn-taking latency.

model release

Unsloth Releases GGUF Quantizations of Meta's Muse Glimmer 30B Agentic Model

Unsloth has published GGUF quantizations of Muse Glimmer-30B, a dense 29.6B-parameter causal transformer with a dedicated perception encoder, attributed to Meta Superintelligence Lab in the model card. The model targets autonomous agentic tasks on consumer hardware with a 131,072-token context window and 4-bit quantization under 20GB.

model release

Meta Releases Muse Glimmer, First Open-Weight Model Since Llama 4, Paired With Zuckerberg Manifesto on Distillation

Meta has released Muse Glimmer, a 30-billion-parameter open-weight model under Apache 2.0, its first open release since Llama 4 in spring 2025. The launch comes with a Zuckerberg essay defending distillation of rival models and hinting at a future 'dynamic auction' pricing scheme for compute.

model release

Meta Open-Sources Muse Spark 1.2, Announces On-Device Model Family Muse Glimmer

Meta CEO Mark Zuckerberg announced the company will open-source its Muse Spark 1.2 model and launch a new on-device model family called Muse Glimmer. The move positions Meta against closed-model rivals OpenAI and Anthropic and against Chinese open-weight labs like DeepSeek and Alibaba.

Comments

Loading...