Tavus says 48% of testers mistook its Griffin video AI for a real person on a one-minute call
Tavus has introduced Griffin, which it calls the first 'Human Interaction Model' for real-time face-to-face video conversation. In a Tavus study, 48% of participants believed Griffin was a real person after a one-minute call, versus a 2% maximum for earlier systems. A limited research preview, Griffin-Lite, is open only to select testers.
Tavus says 48% of participants in its own study believed its new Griffin model was a real person after a one-minute video call. The San Francisco startup claims earlier systems topped out at 2%. Griffin is billed as the first "Human Interaction Model" (HIM), a model class built to understand and conduct face-to-face conversation in real time.
What Griffin does
According to Tavus, Griffin processes speech, facial expressions, tone of voice, gestures, and pauses while it receives and generates video. That makes it an audio-video conversational system rather than a text model with a video layer on top.
The numbers Tavus is reporting
- Human-identification study: 48% of participants believed Griffin was a real person after a one-minute video call. Tavus says previous systems maxed out at 2%.
- Nvidia-run test: In what Tavus describes as an independent Nvidia test measuring how human an AI feels in direct audio-video conversation, Griffin scored 3.83. Real humans scored 3.92, and the previous best AI model scored 2.80.
Both sets of results come from Tavus and have not been independently verified. The scoring scale for the Nvidia test, the participant count, and the study methodology were not given in the coverage of the announcement. Tavus's research report contains further detail.
Availability
A preview called Griffin-Lite is available to "select testers as a research preview." Tavus says a more capable version will follow once "safety concerns are addressed." There is no public API, so the following are not disclosed:
- Pricing: not yet disclosed
- Context window: not yet disclosed
- Parameter count: not yet disclosed
- Training cutoff: not yet disclosed
- General availability date: not yet disclosed
Intended use cases
Tavus lists tutoring, practicing difficult conversations, and camera-based tech support as potential applications.
Company background
Tavus was founded in 2020 and has raised about $64 million. It started with personalized AI video for sales and marketing, then moved into live video conversations with digital personas. Griffin extends that line into a purpose-built interaction model.
What this means
The headline figure is a vendor-run result on a very short interaction. One minute is enough to test first impressions, but it says little about whether a persona holds up over a longer or more adversarial conversation. Without the sample size and protocol, the jump from 2% to 48% can't be assessed, and the Nvidia-linked scores need the same scrutiny until the methodology is published and reproduced.
If the results hold, the practical consequence is on the trust side. A system that a large share of people can't distinguish from a human on a live video call weakens casual visual verification, which matters for fraud, impersonation, and social engineering. Tavus's decision to hold back a more capable version over unresolved safety concerns suggests it sees that risk too.
For builders, Griffin is not yet something to integrate. Until Tavus opens access and publishes pricing, latency figures, and evaluation details, it is a claim to track rather than a tool to adopt.
Related Articles
Nvidia's SoL-Pi Cuts Coding Agent Token Usage by Up to 49% Through Automated Harness Optimization
A new Nvidia research system called SoL-Pi automatically rewrites the control logic of coding agents rather than the underlying model, cutting token usage by up to 49% while keeping performance nearly intact. The approach could shift efficiency gains in AI agents from model-level tricks to harness-level engineering.
Graphite: Opus 5.5 uses 'this matters' 116x more than humans as AI writing tells persist
Marketing firm Graphite identified 13,000 phrases that appear at least twice as often in AI-generated writing as in human writing. Claude Opus 5.5 uses "this matters" 116 times more than humans, while OpenAI's Astra favors "corrective framing" more than 100 times as often. Em-dash use has collapsed across frontier models, but total tells are holding steady, according to Graphite.
Ai2 releases Olmo-core 3, an open MoE training stack benchmarked at 1.2T parameters
Ai2 released Olmo-core 3, an open training framework for large mixture-of-experts models. The company reports about 2.7× the throughput of its earlier FSDP-based implementation and benchmarks up to 1.2 trillion total parameters on 512 NVIDIA B300 GPUs. It is the infrastructure for Ai2's next MoE-based Olmo, not a model release.
Anthropic Red Team: GLM-5.3 Matches Claude on Binary Exploitation for First Time
Anthropic's Frontier Red Team reports that Zhipu AI's GLM-5.3 achieved full control flow hijacks in 4% of binary exploitation trials, versus 6% for Claude Mythos Preview. Predecessor models Claude Opus 4.6 and GLM-5.2 scored zero, marking what Anthropic calls a crossed threshold in offensive cyber capability.
Comments
Loading...