Models

Get clear and helpful insights directly from audio files. Identify speakers, understand the key points they’ve made – and grasp the sentiment behind those points.



Advanced AI transcription capabilities

A unified voice experience that cleans up speech, understands intent, and executes tasks.

Disfluency clean-up

Filters pauses, “ums”, “ahs” and other filler words, to produce polished text with accurate punctuation and useful formatting – at the speed of speech.

Function calling

Delegates complex tasks to other Gemini models—enabling image generation, live search, text summarization, and file analysis.

Adaptable voice editing

Refine your thoughts in the moment – correcting details, clarifying spellings, or change the writing style without missing a beat.



Model information

Name
3.5 Transcribe
Status
Preview
Input
  • Text
  • Audio
Output
  • Text
Input tokens
96k
Output tokens
32k
Availability
  • Google Antigravity
  • Gboard Rambler
  • Gemini API
  • Gemini App
  • Gemini App on MacOS
  • Google AI Studio
  • Google Enterprise Agent Platform
  • Google Workspace
Documentation
View developer docs
Model card
View model card

Try AI transcription