Google announced Gemini 3.8 Live with Live Avatar, a capability that gives enterprise conversational agents a synchronized visual persona. The system combines real-time video generation and speech models to present an on-screen agent that moves and speaks in sync with the conversation.
Live Avatars provide precise lip-syncing, natural facial expressions, and fluid turn-taking designed for customer service and interactive walkthroughs. The visual output is tied to Gemini's speech and reasoning models so the avatar matches the spoken responses and conversational pacing.
A central feature is asynchronous tool calling. The avatar can trigger tool calls and fetch data in the background without pausing the interaction. That enables scenarios where the agent performs tasks such as booking or checking a customer into a hotel while the user continues to talk. The process keeps the conversation flowing rather than shifting the user into wait states.
Google says Live Avatars adapt lip-syncing and facial expressions across languages. The system supports 97 languages and reportedly maintains video quality when switching between them.
Gemini 3.8 Live with Live Avatar launches now but is restricted to Gemini Enterprise subscriptions. Google positions the feature for enterprise customer service and interactive demos rather than as a consumer-facing product.
What this means for users and companies
For enterprises that adopt Live Avatars, the value proposition centers on more natural-feeling agent interactions and fewer interruptions while backend tasks run. For customers, conversations may feel more like speaking with a person because of synchronized mouth movements and expressions. The enterprise-only rollout means customers will see Live Avatars primarily through companies that integrate Gemini Enterprise into their support channels.
The published example shows an avatar checking a user into a hotel while the user continues to converse — demonstrating the combination of visual presence and background tool activity.
Live Avatars are an addition to Gemini's capabilities that pair real-time visual feedback with background task execution and broad multilingual handling. Right now, the feature is a commercial offering for enterprises rather than an end-user product.