Techcrunch iconTechcrunchSep 26, 2026 ~7 min source read

A TechCrunch reporter had a digital avatar made of them — you can talk to it

At Synthesia’s New York office the reporter had photos and a short voice recording turned into a speaking, listening avatar trained to discuss a single TechCrunch story about venture fraud.

I created an interactive digital avatar of myself — and you can talk to it

Share this story

Send the public story page.

Useful takeaways from this story.

Synthesia produced both scripted and interactive versions of a journalist’s digital avatar using photos and a two-minute voice sample.

Deployments fall into three product types: classic scripted video avatars, Sessions for interactive roleplay or surveys, and an API for building custom interactive avatars.

# What happened

A TechCrunch reporter accepted Synthesia's invitation to get a personal digital avatar during the company's New York office opening. The studio captured photos and a two-minute voice recording, then produced several avatars: scripted personal avatars (with and without glasses) and two interactive avatars trained specifically to answer questions about one TechCrunch story on venture-backed startup fraud.

# How the interactive avatar was built

The creation process took a few days. The reporter sat in a mini film studio where team members photographed and recorded them. Synthesia created three deliverables: a personal avatar that reads typed scripts, a personal avatar variant, and two interactive avatars that can listen and respond. The interactive versions were deliberately trained only on the chosen article so they would answer questions related to that story.

# Technology pipeline (in plain terms)

The avatar's behavior is produced by a multi-stage pipeline:

  • A voice-to-text component turns a user's spoken question into text.
  • A language agent interprets that text and determines the correct response or action, constrained to the article's content in this deployment.
  • A text-to-voice component renders the response in the reporter's voice profile.
  • A video model animates the on-screen avatar to match the audio.

# The reporter's experience

# Product lineup and commercial options

Synthesia organizes its offerings into three product areas:

  • A video-creation platform with classic avatars that read supplied scripts.
  • Sessions, an interactive product for roleplay, training and surveys that can score responses and simulate conversational scenarios.
  • An API platform that lets customers combine Synthesia's video and voice models with external services to create custom interactive avatars.

# Practical implications for readers

If you are a content creator, communicator, or company exploring avatar use, this account shows several concrete points:

  • Fidelity varies: scripted avatars may look and sound closer to the person than interactive ones trained narrowly on specific content.
  • Control is possible: avatars can be trained to restrict responses to specific materials, which limits unexpected answers.
  • Deployment choices matter: model and hosting options affect cost, latency and control.

# Bottom line

More context around this story.

Google、リアルタイムで表情豊かに話すAIアバター「Gemini 3.8 Live with Live Avatar」提供開始 画像1枚で独自アバターも
Itmedia iconItmediaSep 25, 2026

Google、リアルタイムで表情豊かに話すAIアバター「Gemini 3.8 Live with Live Avatar」提供開始 画像1枚で独自アバターも

Googleは、リアルタイム音声対話と動画生成アバターを組み合わせた「Gemini 3.8 Live with Live Avatar」をGemini Enterpriseで提供を開始した。写真1枚から自然に応答するアバターを生成し、多言語や非同期ツール実行に対応。顧客対応や店頭端末などの商用展開を見込む。

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app