Google has released Gemini 3.5 Transcribe to general availability, offering real-time speech-to-text transcription across 85 languages with word-level timestamps and speaker diarization. The model represents Google’s push to make speech processing a core part of the Gemini ecosystem.
Capabilities
Gemini 3.5 Transcribe is a specialized speech-to-text model optimized for real-time and batch transcription:
- Real-time transcription — sub-200ms latency for live speech
- 85 languages — covering the vast majority of spoken languages worldwide
- Word-level timestamps — precise timing for each word in the transcription
- Speaker diarization — automatically identifies and labels different speakers
- Punctuation and formatting — intelligent formatting of transcribed text
- Noise robustness — performs well in noisy environments and with multiple speakers
Performance
Google’s benchmarks show Gemini 3.5 Transcribe outperforming previous generation models across multiple metrics:
- Word Error Rate (WER): 4.2% averaged across all supported languages — a significant improvement over 3.0’s 6.1%
- Real-time factor: 0.08x — meaning it processes 12.5 seconds of audio per second of compute time
- Language coverage: 85 languages (up from 32 in 3.0)
- Speaker detection accuracy: 97.3% for up to 6 speakers
Integration with Gemini Ecosystem
The Transcribe model is designed to work as part of larger Gemini-powered applications:
- Gemini API — direct integration for developers building speech-enabled apps
- Vertex AI — enterprise deployment with full compliance and governance tools
- Google Cloud Speech-to-Text — existing users can upgrade to the Gemini-powered model
- Gemini Assistant — powers real-time voice conversations in Google’s consumer products
Use Cases
The combination of real-time performance, broad language support, and high accuracy opens up several applications:
- Live captioning — accessibility for meetings, lectures, and broadcasts
- Customer service — real-time transcription of support calls
- Content creation — automatic transcription of podcasts, interviews, and videos
- Legal and medical — accurate transcription of depositions, depositions, and patient encounters
- Education — transcription of multilingual classroom content
Pricing
Gemini 3.5 Transcribe is priced competitively with existing speech-to-text services:
- Standard — $0.016 per 15 seconds of audio
- Enhanced — $0.024 per 15 seconds with speaker diarization and word timestamps
- Streaming — $0.020 per 15 seconds for real-time transcription
The pricing makes it accessible for both individual developers and enterprise applications.