Google has revealed a new AI audio model that it says offers improved speech recognition and transcription. The company says Gemini 3.5 Transcribe is more adept at text-to-speech than previous models, with greater accuracy and the ability to automatically detect more than 85 languages. It says the model can adapt a user’s unstructured speech into formatted text. You can use it to make changes with voice commands and it also removes filler words from transcribed speech.
Instead of typing out every thought, Gemini 3.5 Transcribe works with how you actually speak:
✅ Seamlessly handles autocorrections
✨ Removes filler words to provide clean, formatted text
🎯 Understands your natural intention and speaking style
🔊 Accurately captures audio in… pic.twitter.com/17YcsDi7Q0– Google (@Google) August 26, 2026
Google says the model can learn custom vocabulary and unique spellings, and is capable of capturing alphanumeric strings such as order numbers and zip codes. Additionally, the company says Gemini 3.5 Transcribe can assign speech to up to three speakers with word-level timestamps based on pre-recorded audio. This could be useful, for example, for transcribing a podcast.
Perhaps unsurprisingly, Gemini 3.5 Transcribe powers Rambler functionality on Android devices such as the Pixel 11 series phones as well as the Gemini app on macOS (where it can work with other Gemini models to perform agent tasks).
Oddly enough, Google says you’ll soon be able to use text-to-speech in any field on a web page in Chrome. For example, you will be able to dictate responses and messages, and use your voice to invite Gemini. Gemini 3.5 Transcribe is also available in Google Antigravity, the company’s agent development platform, and is coming to Search Live, Gemini Live, Docs, Keep and Gmail. Developers will be able to exploit the model via APIs.
Updated, August 26, 2026, 2:05 p.m. ET: Updated to remove some information that Google has not yet publicly announced.
