All posts
News

DeepMind unveils speech AI model Gemini 3.5 Transcribe

DeepMind unveiled Gemini 3.5 Transcribe on August 26. It improves on Chirp 3 with lower WER, faster latency, two APIs, 85+ languages, and speaker diarization for up to 3 speakers.

Aug 27, 2026 5분 읽기

Why accurate speech-to-text matters again

As speech interfaces expand from smartphones to desks and the web, precise text conversion is critical again. Existing speech models still struggle with background noise, multilingual use, and filler-word removal. The Gemini 3.5 Transcribe that DeepMind announced on August 26 repositions this problem at the API layer.

DeepMind unveils Gemini 3.5 Transcribe

DeepMind released Gemini 3.5 Transcribe as its latest speech AI model. Its biggest leap over the previous Chirp 3 model is lower word error rates and faster response times. It is already available as a preview in Google AI Studio and the Gemini Enterprise Agent Platform, and is in use in the macOS Gemini app and Android Gboard Rambler.

Key performance and numbers

Based on Artificial Analysis benchmarks, it records an average streaming WER of 4.0% and 2.6% on offline audio processing. DeepMind says its text-generation latency is up to 70% lower than prior models.

On the multilingual FLEURS benchmark, it shows streaming 5.50% and non-streaming 5.04% WER. It automatically identifies 85+ languages and jointly handles regional accents and code mixing. It also reliably strips special characters and alphanumeric IDs from transcripts. The speaker diarization feature accurately separates up to 3 speakers, and 3 or more is in testing.

Two APIs for different workflows

Gemini 3.5 Transcribe ships as two APIs. gemini-3.5-transcribe-live supports real-time streaming transcription for short-latency interactive apps. gemini-3.5-transcribe processes offline audio files, meetings, and call recordings, offering word-level timestamps and speaker attributes. Both APIs support custom vocabulary so registering domain terms or product names directly improves conversion accuracy.

Real product usage examples

Gboard's Rambler turns spoken thoughts into structured text and automatically removes filler words. Users can highlight text for rewriting or summarize documents. Antigravity captures the user's meeting screen and chat history together and extracts file names and document content precisely.

The macOS Gemini app supports local file transcription, text repositioning between apps, and image generation via function calling. Chrome is also expected to soon support audio input in web input fields.

Developer access and current status

Developers can directly access gemini-3.5-transcribe-live in Build mode in Google AI Studio to build prototypes. Enterprise environments are available through the Gemini Enterprise Agent Platform, and the upcoming Gemini Enterprise for Customer Experience. The function-calling feature is confirmed only in the macOS Gemini app for now. The claim that 3-or-more speaker diarization is in testing should be treated as tentative because actual production rollout timing may shift. Agora, LiveKit, LangChain, and Vercel developer platforms have already announced Live API integrations, so checking their SDKs directly is worthwhile.

Conclusion

Gemini 3.5 Transcribe clearly outperforms Chirp 3 on WER and response speed. Its two APIs—real-time streaming and offline processing—align with development workflows, and support for 85+ languages and custom vocabulary strengthens practical adoption. If you are planning a speech-interface project, first review the preview in Google AI Studio.

#DeepMind#Gemini 3.5 Transcribe#Speech AI#AI News#Google AI Studio
Robeedau

Curated, fact-checked, and edited by a single operator before publishing.