AI News

Google launches Gemini 3.5 Transcribe

Google DeepMind released Gemini 3.5 Transcribe in public preview, a speech-to-text model it says cuts word error rates to 4.0% streaming and 2.6% non-streaming and reaches final transcripts 70% faster than Chirp 3.

Google DeepMind on Tuesday introduced Gemini 3.5 Transcribe, a speech-to-text model now in public preview through the Gemini API in Google AI Studio, Google Antigravity, and the Gemini Enterprise Agent Platform.

The model ships as two endpoints: gemini-3.5-transcribe-live for bidirectional streaming with sub-second latency via the Live API, and gemini-3.5-transcribe for pre-recorded audio via the Interactions API, with speaker attribution and word-level timestamps for up to three speakers (more is experimental).

Google cites Artificial Analysis measurements of an average 4.0% word error rate for streaming and 2.6% for non-streaming, and says time to final transcription improves 70% over its previous model, Chirp 3. On the multilingual FLEURS benchmark across a set of top languages and locales, it reports 5.50% WER streaming and 5.04% non-streaming. The model auto-detects and transcribes more than 85 languages, removes filler words, resolves self-corrections, and adapts to supplied custom vocabulary. In the Gemini macOS app it can also make function calls to hand off image generation or file analysis to other Gemini models.

Distribution is already broad: Rambler in Gboard on Android in select countries and languages, the Gemini app on macOS in English, Build mode in AI Studio, and Chrome "coming soon." Google names Agora, Fishjam, LangChain, LiveKit, Pipecat, Vercel, and Vision Agents as Live API platforms, and Vivo, Intellitek Health, and Lingopal as early users.

Google did not disclose pricing, general-availability timing, or the composition of the Artificial Analysis test set.

All stories