AI

Google bets Gemini 3.5 Transcribe can clean audio the way a human editor would

Adrian Kessler

Gemini 3.5 Transcribe does not just capture what someone says. It converts what they said into what they meant to say. When a speaker corrects themselves mid-sentence — ‘Let’s schedule it for Tuesday, no, Wednesday’ — the model outputs ‘Wednesday’ and discards the correction. Filler words vanish. Background noise does not stop it. The result arrives as polished, formatted text, not raw audio capture.

For anyone who records interviews, podcasts, or multi-language content, this distinction is the entire workflow. Every transcription tool on the market handles ‘what was said.’ The human editor who cleaned it up afterward handled everything else. Google is now building that editor into the model itself.

The model supports two distinct API paths. The Live API handles continuous bidirectional streaming with sub-second latency — designed for real-time voice agents, customer service bots, and anything that needs instant transcription. The Interactions API handles recorded audio: full meetings, call logs, and interviews, returning speaker attribution for up to three participants with timestamps. Both achieve a 4.0% word error rate in streaming mode and 2.6% in non-streaming — measured by independent benchmarking firm Artificial Analysis. Google’s previous speech-to-text model, Chirp 3, required 70% more time to reach final transcription.

The consumer side of the launch is more modest. On Android, Gboard Rambler already uses Gemini 3.5 Transcribe for voice-to-text input. The Gemini macOS app supports it in English via a Speak to Window feature. Chrome is listed as coming soon, which would let users dictate into any web field — replies, posts, forms — without switching apps. No specific date for that rollout has been confirmed.

The model’s limits matter. Multi-speaker attribution caps at three participants; panels and roundtables with larger casts require workarounds. The consumer Rambler app is currently available in ‘select countries and languages,’ not the full 85 the model supports. Google has not announced pricing for the API, which is in public preview through Google AI Studio. Enterprise access comes through the Gemini Enterprise Agent Platform, where pricing is negotiated separately. For smaller creators, the cost question remains open.

The competitive frame is also relevant. OpenAI’s Whisper, still the default for independent developers, achieves word error rates in the 5–7% range on English audio and requires post-processing code to handle filler removal and formatting. Services like AssemblyAI and Deepgram offer speaker diarization but not the built-in natural-language correction Gemini 3.5 Transcribe applies by default. The gap is measurable, though how it holds up across the harder end of the 85-language range — regional accents, low-resource languages, noisy environments — will take independent testing to establish.

The feature that signals the most about Google’s long-term direction is Chrome integration. Once voice input works in any web field, the model is no longer a developer tool — it is infrastructure for how people type. The rollout timeline for that change has not been confirmed.

Tags: , , , , ,

Discussion

There are 0 comments.