Google DeepMind launched Gemini 3.5 Transcribe on August 26 to replace Chirp 3, cutting transcription time by 70% and producing formatted output across 85+ languages.
Ars Technica measured a live-speech error rate of 5.5%, down from 7.32% for Chirp 3; Artificial Analysis found a Word Error Rate of 2.6% for non-streaming. Developers access it through the Live API for real-time streaming capped at 10 minutes, and the Interactions API for pre-recorded audio up to one hour, with timestamps and speaker identification.