# What happened
A TechCrunch reporter accepted Synthesia's invitation to get a personal digital avatar during the company's New York office opening. The studio captured photos and a two-minute voice recording, then produced several avatars: scripted personal avatars (with and without glasses) and two interactive avatars trained specifically to answer questions about one TechCrunch story on venture-backed startup fraud.
# How the interactive avatar was built
The creation process took a few days. The reporter sat in a mini film studio where team members photographed and recorded them. Synthesia created three deliverables: a personal avatar that reads typed scripts, a personal avatar variant, and two interactive avatars that can listen and respond. The interactive versions were deliberately trained only on the chosen article so they would answer questions related to that story.
# Technology pipeline (in plain terms)
The avatar's behavior is produced by a multi-stage pipeline:
- A voice-to-text component turns a user's spoken question into text.
- A language agent interprets that text and determines the correct response or action, constrained to the article's content in this deployment.
- A text-to-voice component renders the response in the reporter's voice profile.
- A video model animates the on-screen avatar to match the audio.
# The reporter's experience
# Product lineup and commercial options
Synthesia organizes its offerings into three product areas:
- A video-creation platform with classic avatars that read supplied scripts.
- Sessions, an interactive product for roleplay, training and surveys that can score responses and simulate conversational scenarios.
- An API platform that lets customers combine Synthesia's video and voice models with external services to create custom interactive avatars.
# Practical implications for readers
If you are a content creator, communicator, or company exploring avatar use, this account shows several concrete points:
- Fidelity varies: scripted avatars may look and sound closer to the person than interactive ones trained narrowly on specific content.
- Control is possible: avatars can be trained to restrict responses to specific materials, which limits unexpected answers.
- Deployment choices matter: model and hosting options affect cost, latency and control.
# Bottom line