// THE VERGE — INTELLIGENZA ARTIFICIALE
Google’s new AI transcription edits out your ‘ums’ and ‘ahs’
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this topic will be added to your daily email digest and your homepage feed.
Posts from this topic will be added to your daily email digest and your homepage feed.
We got a new Gemini Audio model while we’re still waiting for the overdue Gemini 3.5 Pro launch.
We got a new Gemini Audio model while we’re still waiting for the overdue Gemini 3.5 Pro launch.
Posts from this author will be added to your daily email digest and your homepage feed.
Posts from this author will be added to your daily email digest and your homepage feed.
Google has updated Gemini Audio with new transcription capabilities that automatically detect specialized jargon and more than 85 languages. Gemini 3.5 Transcribe is a new addition to the Gemini family that follows the launch of 3.5 Live Translate, and comes as we’re still waiting for Google to release the Gemini 3.5 Pro model that it promised to roll out in June.
Google says that 3.5 Transcribe “represents a major advancement from our previous transcription model, Chirp 3,” especially regarding multilingual performance and wording error rates. The transcription model allows users to “edit naturally with just your voice,” according to Google, and can automatically format text and remove filler words like “um” and “uh.”
Users can provide a customized vocabulary to the model, allowing 3.5 Transcribe to automatically adapt transcription to unique spelling requirements and specialized jargon to prevent those words from being edited manually. It can also attribute speech for up to three speakers in pre-recorded audio, alongside providing word-level timestamps.