Google iconGoogleSep 24, 2026 ~5 min source read

Gemini 3.8 Live with Live Avatar reaches general availability for enterprise agents

Google Cloud has made Gemini 3.8 Live with Live Avatar generally available in Gemini Enterprise, adding synchronized video avatars, native speech-to-speech dialogue, tool calling, and live visual understanding for production voice and multimodal agents.

Power your agents: Gemini 3.8 Live with Live Avatar is now generally available

Share this story

Send the public story page.

Useful takeaways from this story.

Gemini 3.8 Live with Live Avatar is now generally available in Gemini Enterprise with US and EU endpoints, provisioned throughput, and enterprise compliance.

Core capabilities include native speech-to-speech for fluid dialogue, live video avatars with synchronized lip-syncing, background tool and API execution, and live visual understanding of camera feeds and screen shares.

Content provenance and identity safeguards: generated audio and video streams carry SynthID watermarks and custom avatars require enterprise allowlisting and verification.

# What changed Gemini 3.8 Live with Live Avatar is now generally available to Gemini Enterprise customers. The release pairs near real-time video generation with Gemini's live dialogue stack so agents can speak, see, and display a synchronized video persona during conversations. Availability includes US and EU endpoints plus enterprise features such as provisioned throughput, compliance controls, and data governance.

# Core features at a glance

  • Native speech-to-speech: The model handles spoken interaction directly, improving interruption recovery and preserving conversation context and backend transactions while users speak.
  • Live Avatar video: The feature generates animated video avatars with precise lip-syncing and facial expressions that align with the spoken output. Custom avatar creation is gated by an enterprise allowlist and verification.
  • Tool calling: Agents can execute tools and API calls in the background while continuing the conversation so agents can acknowledge requests immediately and keep interacting as operations complete.
  • Live visual understanding: The system can process live camera feeds and screen shares simultaneously with audio input, enabling agents to "see" what users show and update structured artifacts (for example, claim forms) in real time.
  • Multilingual support: The live model understands and speaks 97 languages and can automatically detect language, switching languages mid‑conversation without degrading visual fidelity.

# Trust, provenance, and identity controls

# Demos and integration examples Google published three demos showing core workflows:

  • Interactive custom avatar: Build a custom avatar by adding system instructions, a single reference photo, and an audio sample.
  • Real-time voice agent with ADK and Gemini Live API: Developers can stream real-time audio directly to the Gemini Live API and use ADK to define agents, manage runners, and maintain session memory without a traditional speech-to-text pipeline.

# Customer use cases and pilot feedback Several enterprise customers have piloted or adopted the capability. Cox Automotive (Autotrader) built a shopping assistant that uses live screen-highlighting and tool-calling to guide vehicle discovery through conversation. Equal AI reported improved interruption handling, multilingual conversations, and tool-call reliability for high-volume, multilingual call handling. Salesforce's Agentforce collaboration demonstrates interest in combining real-time multimodal capabilities with agentic workflows.

# Practical deployment notes Start points for engineering teams include the Gemini Live API and Google's Agent Development Kit (ADK). The ADK helps teams define agent behavior, manage session state, and stream live audio to the model without first converting speech to text. Enterprises should plan for compliance review, avatar verification if creating custom personas, and for integrating SynthID tracking into content workflows.

# Bottom line This release focuses on production-ready conversational agents that combine live voice, background tooling, and a synchronized video presence. Enterprises evaluating real-time customer-facing agents should consider the trade-offs between ready-made avatars and custom personas, verify governance and verification workflows, and prototype integrations using ADK and the Gemini Live API.

More context around this story.

AI бЂ”бЂЉбЂєбЂёбЂ•бЂЉбЂ¬бЂЂбЂ­бЂЇ бЂЎбЂ™бЂјбЂ”бЂєбЂ†бЂЇбЂ¶бЂё бЂњбЂ±бЂ·бЂњбЂ¬бЂ”бЂЉбЂєбЂё
Medium iconMediumSep 5, 2026

AI бЂ”бЂЉбЂєбЂёбЂ•бЂЉбЂ¬бЂЂбЂ­бЂЇ бЂЎбЂ™бЂјбЂ”бЂєбЂ†бЂЇбЂ¶бЂё бЂњбЂ±бЂ·бЂњбЂ¬бЂ”бЂЉбЂєбЂё

AI (Artificial Intelligence) နည်းပညာက အá€á€¯á€¡á€á€»á€­á€”်မှာ နေရာá€á€­á€¯á€„်းမှာ ရှိနေပါပြီዠဒါပေမဲ့ “AI ကို ဘယ်ကနေ စလေ့လာရမလဲአအမြန်ဆုံး á€á€á€ºá€™á€¼á€±á€¬á€€á€ºá€¡á€±á€¬á€„်â

Loading more related stories...

Keep reading in the app

Open the app view to save this story, compare related coverage, and continue from the same source.

Open in app