Google has unveiled a new AI audio model, Gemini 3.5 Transcribe, which the company says delivers improved speech recognition and transcription capabilities compared to earlier versions. The model joins Gemini 3.5 Live and Gemini 3.5 Live Experimental in the Gemini Audio family, and Google states it is more precise at speech-to-text, automatically detecting more than 85 languages. For US users who rely on voice input, the model is designed to convert unstructured speech into formatted text, handle edits via voice commands, and strip out filler words.
A key application highlighted by Google is the upcoming ability in Chrome to use speech-to-text in any web field, allowing users to dictate replies, compose posts, and prompt Gemini with their voice directly on web pages. The model can also learn custom vocabulary and unique spellings, and is touted as adept at capturing alphanumeric strings such as order numbers and postal codes. Additionally, for pre-recorded audio like podcasts, Gemini 3.5 Transcribe can attribute speech to up to three distinct speakers with word-level timestamps.
The company says the model understands natural intent and speaking style, and it works seamlessly with self-corrections, meaning users can talk as they normally would without needing to structure their wording. This feature set positions the tool as a more flexible alternative to typing, particularly for drafting longer messages or notes in a browser environment. Google emphasizes that the model is built to match how people actually talk rather than forcing users to adapt to a rigid command structure.
Gemini 3.5 Transcribe is already powering the Rambler feature on Android devices, including the Pixel 11-series phones, and the Gemini app on macOS, where it can collaborate with other Gemini models for agentic tasks. On macOS, this integration allows the transcription model to work alongside other AI capabilities to execute multi-step actions. This cross-platform availability means US consumers on both mobile and desktop operating systems can access the new functionality.
Beyond consumer apps, Google is making the model available in Google Antigravity, its agentic development platform for builders, and it is being integrated into Search Live, Gemini Live, Docs, Keep, and Gmail. Developers can also tap into Gemini 3.5 Transcribe through APIs, opening the door for third-party applications to use the same speech recognition technology. This broad rollout suggests the model will serve both everyday users and enterprise or developer needs across Google鈥檚 ecosystem.
The announcement positions Gemini 3.5 Transcribe as a central piece of Google鈥檚 audio AI lineup, with a focus on practical, real-world speech handling. By combining automatic language detection, speaker attribution, and clean formatting, the model aims to reduce the friction of voice-to-text for both quick inputs and longer media transcriptions. Google has not specified a release date for the Chrome web-field feature, only that it will be available soon. The model is currently live in the named products and platforms, with API access for developers.
More AI news from TechManNews.







