What happened
Google introduced Gemini 3.5 Transcribe, a dedicated speech-to-text model for audio files and real-time streaming. It supports more than 85 languages, language switching, speaker identification, word-level timestamps and custom vocabulary.
Why it matters
Its Smart transcription mode can remove filler words, understand self-corrections and organize speech into a more usable format. This makes transcription a starting point for summaries, CRM updates, tasks, support analysis, captions and content.
Limits
Audio files can be up to one hour. With diarization or word timestamps, the limit is up to 30 minutes per request. Attribution for three or more speakers is experimental. Live sessions are limited to ten minutes.
Google reports about 70 percent faster time to final transcription than Chirp 3, with Word Error Rate figures of 5.50 percent for streaming and 5.04 percent for non-streaming use cases. These are Google-reported measurements and should be tested on an organization’s own recordings.
For legal, medical or sensitive conversations, retain the original verbatim version as well. Smart transcription is more convenient, but it includes interpretation.

