Google Gemini 3.8 Live - Real-Time Voice Model for Multilingual Agents

Google Gemini 3.8 Live is a real-time multilingual voice model for agents that run customer-facing spoken dialogue in many languages. Google DeepMind shipped two versions on September 15, 2026 - the cost-efficient Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, which narrates its reasoning aloud while it works.

Core Features

  • Speech-to-speech architecture: audio goes in and audio comes out with no separate transcription or text-to-speech stage, so it handles interruptions mid-sentence.
  • Language switching: it detects and switches among 97 languages on the fly, even changing mid-call without restarting.
  • Background tool use: the model triggers API and tool calls while the conversation continues, then folds results back into what it is saying.
  • Visual grounding: it reads camera feeds, images, and shared screens in near real time through the Gemini Live API.
  • SynthID watermarking on every generated audio clip, plus a published model card with safety details.

Use Cases

  • Multilingual call centers that serve customers who switch languages mid-conversation.
  • Enterprise voice agents for bookings and support spanning several back-end APIs.
  • Accessibility and hands-free assistants that benefit from low-latency spoken answers.

Pricing

Live API access costs $0.005 per minute of audio input and $0.018 per minute of audio output; Search Live is free and the Gemini app ties the models to AI Pro at $19.99 and Ultra from $99.99 per month. Browse related assistants in our AI chat tools directory.

Our Take

Best for global support lines that must switch languages without breaking the call. The trade-off is that Google publishes no detailed API rate card at launch, and the models are hosted only - there are no open weights to self-host.

FacebookXWhatsAppEmail