Google is pushing the boundaries of speech artificial intelligence with the introduction of Gemini 3.5 Transcribe, hailed as its most precise speech-to-text model to date. This advanced system is designed to move beyond basic dictation, offering enhanced capabilities for handling natural conversations, specialized vocabulary, and multiple speakers, ultimately making transcribed content more functional across Google's ecosystem.
Enhanced Accuracy for Natural Speech
Addressing the complexities of real-world conversations, Gemini 3.5 Transcribe excels at cleaning up common speech patterns. It automatically removes disfluencies and filler words, handles self-corrections made mid-sentence, and formats transcribed text for clarity. A key feature is its ability to recognize custom vocabulary, specialized jargon, and unique spellings, crucial for developers building sophisticated voice agents and audio-based applications.
Broad Language Support and Speaker Identification
The new model supports an extensive range of over 85 languages, including various accents and dialects, ensuring global applicability. For recorded audio, Gemini 3.5 Transcribe can accurately identify up to three distinct speakers and provide word-level timestamps, a significant improvement for meeting summaries or call logs. Experimental support for more than three speakers is also under development.
Performance and Integration
Google highlights impressive performance metrics for Gemini 3.5 Transcribe. Artificial Analysis measurements indicate an average Word Error Rate (WER) of just 4% for streaming applications and 2.6% for non-streaming use cases. Furthermore, Google states that the time to final transcription is 70% faster compared to its previous Chirp 3 model. On the FLEURS benchmark, the model achieved a 5.50% WER in streaming and 5.04% in non-streaming tests.
Developers can access this powerful model through two dedicated APIs:
- Live API: Designed for continuous bidirectional streaming, offering sub-second latency for real-time applications.
- Interactions API: Optimized for recorded audio such as meetings and call logs, providing speaker attribution and timestamps.
Google is also integrating Gemini 3.5 Transcribe's capabilities across its product suite. On Android, Gboard's Rambler now converts speech into formatted text, allowing voice-based corrections and style changes. Google Antigravity leverages screen context and chat history to further improve transcription accuracy. The Gemini macOS app uses voice commands for analyzing local files, repurposing text, and generating images. Additionally, Chrome will soon support dictation directly into web fields.
Availability
Gemini 3.5 Transcribe is currently available in public preview through Google AI Studio and Google Antigravity. Enterprise clients can access its features via the Gemini Enterprise Agent Platform. The Gemini macOS app supports the model in English, while Rambler on Android is accessible in select countries and languages.