Synthesia, a video-generation startup originally founded in the United Kingdom, unveiled an interactive digital replica of a TechCrunch journalist during a visit to its new New York office. The company, which recently announced a valuation of four billion dollars and annual recurring revenue exceeding one hundred million dollars, positions itself alongside rivals such as D-ID, HeyGen and Colossyan. Its head of corporate affairs, Alexandru Voica, presented the avatar as the latest addition to the firm’s public-relations toolkit.
To create the likeness, the journalist entered a compact studio inside the office where technicians captured a series of photographs and recorded a two-minute audio sample. After providing consent, the team generated a basic avatar capable of delivering any supplied script, as well as two interactive versions that can listen and respond. Both the static and interactive models were produced in versions with and without glasses to match the subject’s appearance.
The interactive twin relies on a chain of AI models: a speech-to-text engine transcribes spoken input, an agentic language model interprets the text and decides on actions, a text-to-speech system synthesizes the reply, and Synthesia’s proprietary video model animates the facial movements. Customers may replace any component with alternatives from providers such as Cartesia, ElevenLabs, Google or OpenAI, and they can host the service on their own cloud infrastructure or use Synthesia’s managed hosting.
Synthesia organizes its offerings into three segments. The first is a video-creation platform where users type a script and a classic avatar lip-syncs the content. The second, called Sessions, provides an agentic environment for role-playing scenarios such as sales pitches, scoring user performance and allowing real-time interaction. The third is an API suite that lets developers combine Synthesia’s video and voice models with external services to build custom interactive experiences.
The journalist’s avatar was trained exclusively on a single investigative piece about venture-backed startups and fraud, making its responses deterministic: any query outside that topic was redirected back to the story. When asked about personal history, the twin repeatedly cited the article instead of providing new information. Observers noted that the synthetic voice did not perfectly match the original timbre, while the subject’s mother described the result as “amazing,” according to the TechCrunch report.
Industry analysts see such avatars as a potential tool for corporate communications, training and even media presentation, but they also raise concerns about authenticity and audience trust. The deterministic design used in this demonstration limits the risk of unsupervised dialogue, yet future nondeterministic versions could generate unpredictable content. Synthesia’s rapid prototyping,delivering the journalist’s twins within a few days,highlights how quickly personalized AI representations can be deployed across enterprises.