Synthesia created an interactive digital double of a TechCrunch journalist from a photo session and a two-minute voice recording. Unlike a conventional avatar that reads a script, the result could listen to questions and answer them, although this test version was deliberately limited to discussing one previously published article.
The system combines speech recognition, a language model, text-to-speech and Synthesia’s video model. Customers can use alternative voice or language services and choose where the avatar is hosted. The company already sells scripted video tools and Roleplay Sessions, which lets employees practice tasks such as sales pitches with an avatar that responds and scores them.
The experiment also showed the limits of a synthetic stand-in. The avatar redirected personal questions rather than improvising, and friends thought its voice and appearance were close but not fully convincing. That narrow behavior reduced the risk of fabricated personal answers, but it also made clear that likeness is easier to reproduce than trust. For workplaces considering digital representatives, consent, strict topic boundaries and clear disclosure remain as important as visual realism.