Google announces Gemini 3.5 Transcribe for AI-powered speech-to-text
热度趋势
趋势数据积累中
百分比基于当前可用热度信号,而非评论数或独立用户人数。
Gemini 相关模型动态已经出现,适合跟踪能力变化、生态影响和后续可用性。
谷歌发布了Gemini 3.5 Transcribe,这是一款旨在简化语音输入的AI模型。它通过编辑掉“嗯”和更正,输出精炼的AI文本。该模型已为Pixel 11上的Gboard“Rambler”功能提供支持,并将很快在整个谷歌生态系统中推出。尽管该模型在清理不一致和口头结巴方面表现良好,但其缺点在于AI会改变用户所说的措辞,这可能不适用于所有情况。
While we wait (possibly in vain) for Gemini 3.5 Pro to launch, Google is releasing a different model in the 3.5 branch. The company has announced Gemini 3.5 Transcribe, an AI model designed to streamline voice input by editing out “ums” and corrections, outputting polished AI text. This model already powers the Gboard “Rambler” feature on the Pixel 11, but it’s about to appear throughout the Google ecosystem.
According to Google, Gemini 3.5 Transcribe is much faster and more accurate than its previous voice-to-text engine, known as Chirp 3. The new AI model should be about 70 percent faster from voice to final transcribed text, and the live-speech error rate has dropped to 5.5 percent. That’s only a little better than Chirp 3, which Google measures at 7.32 percent. Still, it’s a pain to fix typos when you’re using voice input, so any improvement here is beneficial.
Credit: Google
Credit:
The new model isn’t just better at hearing the words you say; it’s also supposedly better at getting to the heart of what you meant to say. As you speak, Gemini 3.5 Transcribe can remove the awkward “ums” and “uhs” that clutter your stream of consciousness. It can also edit text on the fly if you need to correct yourself and refer to your provided custom vocabulary to handle “specialized jargon.” All this works in 85 languages and with up to three speakers in pre-recorded audio.
The drawback, of course, is that you’re relying on the AI to accurately get the gist of your speech. For short blocks of text, the model seems good at cleaning up inconsistencies and verbal stumbles (based on my testing with Rambler), but the AI does technically change the wording of what you said, and that may not be appropriate for all situations.