Gemini Audio

Live dialogue

Fluid and natural live dialogue and translation capabilities, for powerful voice-first applications

Models

Create agents capable of handling complex tasks and using tools, while engaging in natural conversations.



Build voice agents

See how you can use 3.8 Live Extended Thinking to build voice agents.


Live speech translation

Overcomes language barriers by using Gemini’s speech-to-speech translation capabilities.

Broad language coverage

Delivers fluid speech-to-speech translation across 70+ languages and 2,000 language pairs.

Consistent intonation

Preserves the speaker’s original intonation, pacing and pitch to capture not just what they said, but how they said it.

Multilingual input

Translates multiple languages in a single session – no need to change the settings.

Automatic language detection

Identifies the language being spoken and begins translation, without being told what it is.

Low latency

Minimizes processing lag to eliminate awkward pauses – and keep conversations flowing naturally.



Model information

Name3.8 Live Extended Thinking3.8 Live3.5 Live Translate
StatusGeneral AvailabilityGeneral AvailabilityGeneral Availability
Input
  • Text
  • Image
  • Video
  • Audio
  • Text
  • Image
  • Video
  • Audio
  • Audio
Output
  • Text
  • Audio
  • Text
  • Audio
  • Text
  • Audio
Input tokens128k128k128k
Output tokens64k64k64k
Availability
  • Gemini App
  • Google AI Studio
  • Gemini API
  • Google Enterprise Agent Platform
  • Google Search Live
  • Gemini App
  • Google AI Studio
  • Gemini API
  • Google Enterprise Agent Platform
  • Google Workspace
  • Google Translate
  • Google AI Studio
  • Gemini API
DocumentationView developer docsView developer docsView developer docs
Model cardView model cardView model cardView model card

Try Live dialogue