Google DeepMind released Gemini 3.5 Transcribe, a speech-to-text model the company bills as its most precise yet, into public preview for developers and enterprises.
Per Google DeepMind, the model converts raw audio directly into clean, formatted text, smoothing over self-corrections, filler words, and the awkward pauses that trip up conventional speech recognition. It also recognizes custom vocabulary, adapting transcriptions to specialized jargon and unique spellings.
Developers get two entry points. The Live API serves real-time, bidirectional streaming with sub-second latency through the gemini-3.5-transcribe-live model for interactive voice apps, while the Interactions API processes recorded audio, meetings, and call logs, returning speaker attribution with word-level timestamps.
Independent measurement by Artificial Analysis puts the average word error rate at 4.0 percent for streaming and 2.6 percent for non-streaming use, with FLEURS benchmark scores of 5.50 and 5.04 percent across a set of top languages.
Latency is the other gain: time to final transcription improves by 70 percent over Google's previous transcription model, and the model auto-detects more than 85 languages with regional accents while attributing speech from up to three speakers in pre-recorded audio. Support for more than three speakers remains experimental.
Beyond accuracy, the model handles the messiness of real speech. It absorbs self-corrections such as "let's meet Tuesday, no, Wednesday", strips filler words, and auto-formats the output. A function-calling feature delegates heavier jobs, such as image generation and file analysis, to other Gemini models, and that capability is live in the Gemini app on macOS.
Google is also wiring 3.5 Transcribe into everyday surfaces. On Gboard for Android, the new Rambler feature turns spoken thoughts into formatted text with voice-driven edits; Google Antigravity pairs screen context and chat history, with the user's permission, to transcribe file names and active documents accurately; and Chrome will soon accept dictated replies in any web field.
For developers, the model is in public preview through the Gemini API in Google AI Studio and Google Antigravity, and platforms such as Agora, LiveKit, Pipecat, and Vercel offer it for voice-driven interfaces. For enterprises, it is in public preview through the Gemini Enterprise Agent Platform. A Chrome talk-to-type rollout and wider Android availability are next.













