product update

Mira Murati's Thinking Machines demos real-time 'interaction models' that process audio, video, and text simultaneously

TL;DR

Thinking Machines, founded by former OpenAI CTO Mira Murati, announced it's developing 'interaction models' that continuously process audio, video, and text in real time. Unlike current models that wait for complete input before responding, these models aim to enable simultaneous perception and generation across modalities.

2 min read
0

Mira Murati's Thinking Machines demos real-time 'interaction models'

Thinking Machines, the AI company founded by former OpenAI CTO Mira Murati, announced Monday it's developing "interaction models" that continuously process audio, video, and text while responding in real time.

The company describes current AI models as operating in a "single thread" where they must wait for complete user input before responding, with perception frozen during generation. According to Thinking Machines, this creates a "bandwidth bottleneck" that limits collaboration between humans and AI.

Interaction models aim to solve this by processing multiple modalities simultaneously. The company demonstrated examples including listening for animal mentions in stories, real-time speech translation, and posture detection alerting users when they slouch.

Technical approach

Thinking Machines frames the problem as fundamentally architectural. Current models experience "no perception of what the user is doing or how the user is doing it" until input is complete, and "perception freezes" during generation until the model finishes or is interrupted.

The company's approach enables models to "meet humans where they are, rather than forcing humans to contort themselves to AI interfaces," according to its announcement.

Availability

No technical specifications, benchmark scores, or pricing have been disclosed. Thinking Machines plans to open a "limited research preview" in the "coming months" with a wider release later in 2026. The models are not currently available for testing.

Company background

Murati founded Thinking Machines in February 2025 after departing OpenAI, where she served as CTO. The startup has experienced significant staff turnover, with key members leaving for Meta and returning to OpenAI.

What this means

Thinking Machines is attempting to solve a genuine limitation in current LLM architecture: the turn-based nature of inference. If successful, simultaneous multimodal processing could enable more natural human-AI interaction, particularly for collaborative tasks. However, without technical details on latency, compute requirements, or quality degradation during simultaneous processing, it's unclear whether the approach represents a fundamental architectural advance or an engineering optimization. The company's staff retention issues and lack of concrete deployment timeline suggest early-stage development.

Related Articles

product update

OpenRouter Launches Auto Router Beta: Task-Aware Model Routing Based on Community Spend

OpenRouter has released Auto Router Beta, a task-aware routing system that classifies incoming requests and automatically routes them to popular models based on community spending patterns. The router allows users to filter selections by cost-quality tradeoff preferences.

product update

OpenAI restores chat sidebar in Mac app after user backlash over confusing redesign

OpenAI has updated its ChatGPT Mac app to restore direct access to chat conversations through a prominent sidebar toggle. The fix addresses user complaints following a July 10 redesign that replaced the native Mac client with an Electron-based app and buried the standard chat interface behind Work and Codex features.

product update

NVIDIA NeMo Automodel integrates with Hugging Face Diffusers for distributed video and image model fine-tuning

NVIDIA and Hugging Face have integrated NeMo Automodel with the Diffusers library, enabling distributed fine-tuning of video and image diffusion models without checkpoint conversion. The integration supports models including FLUX.1-dev (12B), Wan 2.1 (1.3B/14B), and HunyuanVideo (13B) with full fine-tuning and LoRA options.

product update

AWS launches Managed Knowledge Base for Bedrock with 6 enterprise connectors and automatic ACL enforcement

Amazon Web Services launched Managed Knowledge Base for Bedrock in general availability, offering a fully managed retrieval solution with six native enterprise connectors including SharePoint, Confluence, and Google Drive. The service handles document parsing up to 500 MB for PDFs, 2 GB for audio, and 10 GB for video, with real-time access control list verification at query time.

Comments

Loading...