Back
Hhackernews·k9294
46
·4 hr ago·Official · Official API

Gemini-3.5-Transcribe

View original
Official announcementGeminiModel release

Heat trend

Collecting trend data

The percentage is based on available heat signal, not comment count or independent people.

Why it matters

An official release brings Gemini model updates — worth tracking for capability changes, ecosystem impact, and follow-up.

AI summary

Gemini 3.5 Transcribe, launched on August 26, 2026, marks a significant improvement over the previous Chirp 3 transcription model.…

Aug 26, 2026

|

Our latest speech-to-text model designed for precise and intelligent real-time transcription.

Diego Melendo Casado

Senior Director, Engineering, Gemini Audio

Luke Leonhard

Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team

Your browser does not support the audio element.

Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.

Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform .

We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:

- Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using gemini-3.5-transcribe-live.

- Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.

Get more precise and intelligent transcription

Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.

- Smart transcription: Seamlessly handles self-corrections (like "let’s meet Tuesday—no, Wednesday" ), removes filler words (“ums” and ‘“ahs"), auto-formats your text.

- Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app .

- More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.

- Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.

- Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.

- Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).

Gemini-3.5-Transcribe · BuzzRadr