Wire flash
Tavus AI video call study: 45% of viewers believed AI was human, while 28% doubted real humans
Editorial responsibility
- No named human review is recorded for this page.
- Source reporting is collected, normalized, translated or condensed automatically when needed.
- Automatically published source-backed update
A study by Tavus found that 120 people watched a 90-second video call and 45% believed the person on screen was real, but it was an AI. In a separate test, 120 people watched two real humans talk, and 28% thought one was a machine. The company gave early access to its smaller model, Griffin Lite. The author notes that looking human on video has become a timing problem, as humans leave roughly 200-millisecond gaps between turns, while turn-based voice AI often gives itself away by pausing. Tavus describes the model as full duplex, meaning it watches and listens while talking, can be interrupted, and notices people entering the frame. NVIDIA researchers created a benchmark from 237 real two-person video call clips, where real humans score 3.92 out of 5. Tavus reports its model scores 3.83, compared to 2.80 for the best published system. The model responds in about 1.9 seconds, versus 2.8 seconds for the other system, while humans answer in 0.9 seconds. The author concludes that language models learned what to say, and face-to-face AI is learning when to say it.
Source report
In a study conducted by Tavus, 120 participants watched a 90-second video call and 45% believed the person on screen was real — but it was an AI. In a separate test, another 120 participants watched a conversation between two real humans, and 28% incorrectly identified one of them as a machine.
The company granted me early access to its smaller test model, Griffin Lite. The 28% figure is the one that stands out: more than one in four viewers doubted that two real people were human.
The Timing Problem
My assessment is that appearing human on video has become a matter of timing. People typically leave gaps of roughly 200 milliseconds between conversational turns, a figure commonly cited from a 2009 study of ten languages. Turn-based voice AI waits for silence before responding, and that pause is often what reveals its artificial nature.
Tavus describes its model as "full duplex," meaning it watches and listens while speaking. The company says the model can be interrupted, cut in on rambling when asked, and notice someone walking into the frame behind you.
Benchmark Performance
NVIDIA researchers developed a benchmark using 237 clips of real two-person video calls. On its generation track, real humans score 3.92 out of 5. Tavus reports its model scores 3.83, compared to 2.80 for the best system in the published paper. The model also delivers responses in approximately 1.9 seconds, while that system took 2.8 seconds.
For context, humans answer in 0.9 seconds on the same test. Nearly a full second still separates the model from a person, and 45% remains below half.
The Core Insight
Language models learned what to say. Face-to-face AI is now learning when to say it.
Source
aakashguptaNeutral / independent