返回
Hhackernews·k9294
46
·5小时前·官方发布 · 官方 API

Gemini-3.5-Transcribe

查看原文
官方公告Gemini模型发布

热度趋势

趋势数据积累中

百分比基于当前可用热度信号,而非评论数或独立用户人数。

推荐理由

官方发布带来Gemini 模型更新信号,适合跟踪能力变化、生态影响和后续落地。

AI 摘要

Gemini 3.5 Transcribe 于 2026 年 8 月 26 日推出,与之前的 Chirp 3 转录模型相比,实现了重大进步。它提供了新功能、改进的词错误率和显著优化的延迟,最终转录时间缩短了 70%。…

Aug 26, 2026

|

Our latest speech-to-text model designed for precise and intelligent real-time transcription.

Diego Melendo Casado

Senior Director, Engineering, Gemini Audio

Luke Leonhard

Chief of Staff, Gemini Audio, on behalf of Gemini Audio Team

Your browser does not support the audio element.

Listen to article

[[duration]] minutes

This content is generated by Google AI. Generative AI is experimental

Today, we’re introducing Gemini 3.5 Transcribe, our most precise speech-to-text model yet, designed for intelligent voice interactions. Unlike conventional speech recognition models that struggle with background noise, complex jargon, and disfluency cleanup, Gemini 3.5 Transcribe converts raw audio directly into accurate, polished, formatted text.

Across our products like the Gemini app and on Android, we’ve seen consumers already benefiting from this transcription model with new voice capabilities like Rambler on Android and in the Gemini app on macOS. Now, developers can build similar capabilities with Gemini 3.5 Transcribe in the Gemini API in Google AI Studio and Gemini Enterprise Agent Platform .

We've built 3.5 Transcribe to plug seamlessly into your developer workflows, whether you’re building voice agents, real-time captioning tools, or post-call analytics pipelines. The model is available across two separate APIs:

- Real-time streaming: Delivers continuous, bidirectional streaming with sub-second latency for interactive voice apps via the Live API using gemini-3.5-transcribe-live.

- Pre-recorded audio processing: Transcribes recorded audio, meetings, call logs, and more with speaker attribution and word-level timestamps via the Interactions API using gemini-3.5-transcribe.

Get more precise and intelligent transcription

Gemini 3.5 Transcribe is designed to capture your natural speaking style to better understand your intent and recognize custom vocabulary, so you can execute tasks with your voice.

- Smart transcription: Seamlessly handles self-corrections (like "let’s meet Tuesday—no, Wednesday" ), removes filler words (“ums” and ‘“ahs"), auto-formats your text.

- Function calling: The model can delegate complex tasks (such as image generation and file analysis) to other Gemini models via function calls. Currently available in the Gemini macOS app .

- More precise transcription: As measured by Artificial Analysis, achieves an average Word Error Rate (WER) of 4.0% for streaming and 2.6% for non-streaming use-cases. It shows strong performance across noisy, real-world environments, accurately capturing alphanumeric entities like postal codes and order IDs.

- Custom vocabulary: Recognizes specialized jargon and unique spellings by seamlessly adapting transcriptions to your provided custom vocabulary.

- Global language support: Automatically detects and transcribes over 85 languages, seamlessly handling regional accents and diverse dialects.

- Multi-speaker identification: Accurately attributes speech in pre-recorded audio with timestamps for up to three speakers (support for 3+ speakers is experimental).

Gemini-3.5-Transcribe · BuzzRadr