Google Gemini 3.5 Transcribe Goes GA — 85 Languages, Real-Time Speech-to-Text

Google releases Gemini 3.5 Transcribe to general availability with real-time speech-to-text across 85 languages, word-level timestamps, and speaker diarization.

Wednesday August 26, 2026 Source: Google
TL;DR — Quick Answer

Google released Gemini 3.5 Transcribe to general availability — real-time speech-to-text across 85 languages with word-level timestamps and speaker diarization. Word error rate improved to 4.2% average (from 6.1% in 3.0), coverage expanded from 32 to 85 languages, and it processes 12.5 seconds of audio per second of compute.

Key Takeaways

Google Gemini 3.5 Transcribe Goes GA — 85 Languages, Real-Time Speech-to-Text — AI news article illustration

Google has released Gemini 3.5 Transcribe to general availability, offering real-time speech-to-text transcription across 85 languages with word-level timestamps and speaker diarization. The model represents Google’s push to make speech processing a core part of the Gemini ecosystem.

Capabilities

Gemini 3.5 Transcribe is a specialized speech-to-text model optimized for real-time and batch transcription:

Performance

Google’s benchmarks show Gemini 3.5 Transcribe outperforming previous generation models across multiple metrics:

Integration with Gemini Ecosystem

The Transcribe model is designed to work as part of larger Gemini-powered applications:

Use Cases

The combination of real-time performance, broad language support, and high accuracy opens up several applications:

Pricing

Gemini 3.5 Transcribe is priced competitively with existing speech-to-text services:

The pricing makes it accessible for both individual developers and enterprise applications.

Frequently Asked Questions

What is Gemini 3.5 Transcribe?

Gemini 3.5 Transcribe is Google's real-time speech-to-text model, released to general availability August 26, 2026. It transcribes 85 languages with word-level timestamps and automatic speaker identification.

How accurate is Gemini 3.5 Transcribe?

The model achieves a 4.2% average word error rate across all supported languages — down from 6.1% in the previous generation — and detects speakers with 97.3% accuracy for up to 6 speakers.

How much does Google's transcription API cost?

Standard transcription costs $0.016 per 15 seconds of audio, enhanced transcription with speaker diarization and word timestamps costs $0.024 per 15 seconds, and real-time streaming costs $0.020 per 15 seconds.

This article is based on the official announcement from Google . Read the original for full technical details.

Related Articles

Back to all news