
Google Launches Gemini 3.5 Transcribe, Its Most Precise Speech-to-Text Model Yet
Google launched Gemini 3.5 Transcribe on August 26, 2026, a new speech-to-text model supporting 85+ languages. It converts raw audio directly into formatted text with automatic filler-word removal, voice-driven editing, and customizable vocabulary for specialized jargon. It also attributes up to three speakers with word-level timestamps. The model succeeds Chirp 3 with improved multilingual accuracy. Developer access is live via the Gemini API.
Published