Google made Gemini 3.8 Live with Live Avatar generally available on September 24, 2026, bringing real-time video presence to enterprise AI voice agents. What was previously a live voice-only capability now includes a visual AI persona that listens, speaks, and lip-syncs in near real time — across 97 languages.
The release lands squarely in the enterprise voice AI market, and it’s a significant signal about where the technology is headed: from audio-only AI agents toward something that looks and behaves more like a video call with a knowledgeable colleague.
What Gemini 3.8 Live Avatar Actually Does
The core capability combines Google’s native speech-to-speech dialogue with near real-time video generation. The result is an AI agent that can hold a fluid conversation while projecting a dynamic visual presence — including lip-syncing, natural facial expressions, and smooth interruption handling.
Key features in the GA release:
- Live Avatar video presence — Animated AI persona that generates video in sync with speech, available across all 97 supported languages
- Live visual understanding — Processes live camera feeds and screen shares alongside audio simultaneously, so the agent can see what the user sees
- Background tool calling — Executes API calls and system actions while the conversation continues, without introducing awkward pauses
- Custom avatars — Enterprise customers can build an avatar from a reference photo via an allowlist program; ready-made avatars are available to all Gemini Enterprise customers immediately
- SynthID watermarks — All generated audio and video carry imperceptible watermarks, ensuring AI-generated content is verifiable and transparent
It’s available through US and EU endpoints with provisioned throughput, enterprise compliance standards, and strict data governance — the table stakes for large organisations with regulated data environments.
Why Enterprises Are Paying Attention
The pricing structure tells you something about the intended market. Google is billing at $1.00 per million video output tokens at introductory rates until 31 December 2026, rising to $7.50 per million at standard rates in 2027. That’s not a developer-tier toy. This is priced for production workloads.
The target use cases Google is highlighting reflect that: customer service deployments at scale, interactive product walkthroughs, virtual support agents that interact naturally with users rather than waiting for typed commands. The YouTube and NFL Sunday Ticket results Accenture published earlier this month (11% sentiment improvement, 37% reduction in handle time) give a sense of what video-capable agents can achieve in high-pressure, high-volume customer moments.
The 97-language capability is significant for global enterprises. An AI voice agent that switches languages mid-conversation based on user input — with lip-syncing that matches — removes one of the biggest practical barriers to deploying voice AI across international customer bases.
What This Means for Business
The floor for enterprise voice AI just rose. Until now, most enterprise voice AI deployments were audio-only — functional, but limited to contexts where a visual presence didn’t matter. Gemini 3.8 Live Avatar makes video-capable AI agents a production reality, not a roadmap item. If you’re evaluating voice AI vendors in 2026, the question is no longer whether video presence is possible but whether your chosen platform can deliver it at enterprise compliance standards.
Customer-facing AI is moving from functional to natural. There’s a meaningful difference between an AI agent that answers questions and one that holds a genuine conversation while maintaining visual presence. The latter changes the user experience fundamentally — particularly for support, onboarding, and advisory interactions where trust matters. This release is Google betting that enterprises will pay for that difference.
The voice AI market is consolidating around multimodal capability. Google’s move here is consistent with a broader pattern: the AI voice market is converging toward platforms that can handle audio, video, vision, and structured tool-calling in a single agent architecture. Standalone voice-only platforms are under increasing pressure to expand, and the companies building multimodal AI employees from the start have an architectural advantage.
For businesses considering voice AI employees for customer service, internal knowledge management, or operational support — the capabilities Google is making enterprise-grade today are the baseline your users will expect within twelve months.
Enterprise DNA builds voice AI employees for businesses through Omni Voice — including customer-facing agents and internal knowledge systems. Talk to us about what a voice AI employee could do for your business.
Your guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideYour guide is ready
Check your downloads folder. If it did not open automatically, use the button below.
Download the GuideSource
Google Cloud BlogTalk it through
Talk it through with Sam
30 minutes on what a Command Centre would look like for your business.
Book a call