Google Launches Gemini 3.5 Transcribe, Its Most Accurate Speech-to-Text Model Yet

Google Launches Gemini 3.5 Transcribe, Its Most Accurate Speech-to-Text Model Yet

Google launched Gemini 3.5 Transcribe on August 26, 2026, a speech-to-text model covering 85+ languages. It converts audio directly into formatted text, removes filler words, supports voice-driven editing, and learns custom vocabulary for specialized jargon. It tracks up to three speakers with word-level timestamps. It succeeds Chirp 3 with better multilingual accuracy. Developer access is live via the Gemini API.

Published

Read at another depth