Gemini 3.5 Audio (Live Translate, Transcribe, Transcribe Live)
Model Cards are intended to provide essential information on Gemini models, including known limitations, mitigation approaches, and safety performance. Model cards may be updated from time to time; for example, to include updated evaluations as the model is improved or revised. See the Google DeepMind site for a comprehensive list of model cards.
Published: August 2026
Model Information
Description
Gemini 3.5 Audio (Live Translate, Transcribe, Transcribe Live) is an addition to the Gemini 3 series of highly-capable, natively multimodal, reasoning models. The models are cost-efficient and fast, optimized for high-volume, latency-sensitive tasks like dialogue and translation. This model card describes the native audio capabilities as additional outputs of Gemini. Information specific to these modalities is specified in-line and referred to as Gemini 3.5 Live Translate, Gemini 3.5 Transcribe, and Gemini 3.5 Transcribe Live referred to collectively as Gemini 3.5 Audio.
Model dependencies
Gemini 3.5 Audio models are based on Gemini 3 Pro.
Inputs
- Gemini 3.5 Live Translate: Audio and text with a token context window of up to 128K.
- Gemini 3.5 Transcribe and Gemini 3.5 Transcribe Live: Audio and text with a token context window of up to 96K.
Outputs
- Gemini 3.5 Live Translate: Audio and text, with 64K token output.
- Gemini 3.5 Transcribe and Gemini 3.5 Transcribe Live: Text, with 32K token output.
Architecture
Gemini 3.5 Audio models are based on Gemini 3 Pro. For more information about the model architecture for Gemini 3.5 Live Translate, Transcribe and Transcribe Live, see the Gemini 3 Pro model card.
Model Data
Training Dataset
Gemini 3.5 Live Translate, Transcribe, and Transcribe Live are based on Gemini 3 Pro. For more information about the training dataset for Gemini 3.5 Live Translate, see the Gemini 3 Pro model card.
Training Data Processing
For more information about the training data processing for Gemini 3.5 Live Translate, Transcribe, and Transcribe Live, see the Gemini 3 Pro model card.
Implementation and Sustainability
Hardware
Gemini 3.5 Live Translate, Transcribe, and Transcribe Live are based on Gemini 3 Pro. For more information about the hardware for Gemini 3 Pro and our continued commitment to operate sustainably, see the Gemini 3 Pro model card.
Software
Gemini 3.5 Live Translate, Transcribe, and Transcribe Live are based on Gemini 3 Pro. For more information about the software for Gemini 3 Pro, see the Gemini 3 Pro model card.
Distribution
Gemini 3.5 Live Translate is distributed in the following channels; respective documentation shared in line:
Gemini 3.5 Transcribe is distributed in the following channels; respective documentation shared in line:
Gemini 3.5 Transcribe Live is distributed in the following channels; respective documentation shared in line:
- Gemini API
- Gemini App
- Google AI Studio
- Google Cloud / Vertex AI
- Google Search Live
- Google Workspace (Gmail, Docs, and Keep)
Our models are available to downstream providers via an application program interface (API) and subject to relevant terms of use. There is no required hardware or software to use the model. For AI Studio and Gemini API, see the Gemini API Additional Terms of Service; for Gemini Enterprise Agent Platform, see Google Cloud Platform Terms of Service. For more information, see Gemini Model API instructions and Gemini API quickstart.
Evaluation
Approach
Gemini 3.5 Audio (Live Translate) was evaluated across a range of benchmarks. See Evals & Methodology for 3.5 Live Translate for more details.
Gemini 3.5 Audio (Transcribe & Transcribe Live) was evaluated across a range of benchmarks. See Evals & Methodology for 3.5 Live Transcribe for more details.
Intended Usage and Limitations
Benefit and Intended Usage
- Gemini 3.5 Live Translate enables low-latency, real-time translation interactions. It processes continuous streams of audio to deliver immediate, spoken responses, creating a natural translation experience for your users.
- Gemini 3.5 Transcribe and Gemini 3.5 Transcribe Live are well-suited for users, developers, and enterprises. Some use cases include: real-time context-aware transcription, agentic voice features, and low-latency conversational agents.
Known Limitations
- Gemini 3.5 Live Translate
- For more information about the known limitations for Gemini 3 Pro, see the Gemini 3 Pro model card.
- Gemini 3.5 Live Translate may exhibit some of the following limitations. Voices can be inconsistent, and voices may shift after long pauses or get stuck on one voice during rapid multi-speaker sessions. Language detection can struggle with non-native accents, similar languages, or rapid language switches. Gemini 3.5 Live Translate is designed to filter out background noise, but not all background audio may be ignored. When set to echo the target language, background noise may introduce artifacts in the translated audio when input audio is in the target language.
- Gemini 3.5 Transcribe and Gemini 3.5 Transcribe Live
- For more information about the known limitations for Gemini 3 Pro, see the Gemini 3 Pro model card.
- Gemini 3.5 Transcribe and Gemini 3.5 Transcribe Live may exhibit some of the general limitations of foundation models, such as hallucinations. There may also be occasional slowness or timeout issues.
- The knowledge cutoff date is January 2025.
Acceptable Usage
For more information about the acceptable usage for Gemini 3.5 Live Translate, Transcribe and Transcribe Live, see the Gemini 3 Pro model card.
Ethics and Content Safety
Evaluation Approach
Gemini 3.5 Audio was developed in partnership with internal safety and responsibility teams. A range of evaluations and red teaming activities were conducted to help improve the model and inform decision-making. These evaluations and activities align with Google's AI Principles and responsible AI approach, as well as Google's Generative AI policies (e.g., Gen AI Prohibited Use Policy and the Gemini API Additional Terms of Service).
Evaluation types included but were not limited to:
- Training/Development Evaluations including automated and human evaluations carried out continuously throughout and after the model’s training, to monitor its progress and performance;
- Human Evaluations conducted by specialist teams across the policies and desiderata to ensure the model adheres to safety policies and desired outcomes;
- Ethics & Safety Reviews conducted ahead of the model’s release.
Additional information: For more information about the evaluation approach for Gemini 3.5 Live Translate, Transcribe, and Transcribe Live, see the Gemini 3 Pro model card.
Safety Policies
For more information about the safety policies for Gemini 3.5 Live Translate, Transcribe, and Transcribe Live, see the Gemini 3 Pro model card.
Frontier Safety Assessment
Gemini 3.5 Audio is part of the Gemini 3 series of models. We evaluated Gemini 3.1 Pro and Gemini 3.7 Flash for Frontier Safety as they are the most generally capable models as of publication of this model card, and they did not reach any Tracked or Critical Capability Levels (T/CCLs) outlined in our Frontier Safety Framework. Our assessments have shown that none of Gemini 3.5 Translate, Gemini 3.5 Transcribe, or Gemini 3.5 Transcribe Live have meaningful new T/CCL-relevant capabilities or material increases in performance with respect to Frontier Safety compared to Gemini 3.1 Pro and Gemini 3.7 Flash, therefore we are confident that none of them are likely to reach any CCLs.
For more information on our Frontier Safety Assessment, read the Gemini 3.1 Pro Model Card and the Gemini 3.7 Flash Model Card.
Risks and Mitigations
For more information about the risks and mitigations for Gemini 3.5 Live Translate, Transcribe and Transcribe Live, see the Gemini 3 Pro model card.