Google Gemini 3.5 Transcribe Enhances Speech-to-Text Accuracy
Google unveiled Gemini 3.5 Transcribe, an AI model enhancing voice-to-text capabilities by removing verbal tics and errors. The technology will soon appear across Google's ecosystem, promising faster, more accurate transcriptions.

Google has officially announced Gemini 3.5 Transcribe, a new artificial intelligence model designed to significantly improve the accuracy and efficiency of speech-to-text conversion. This advanced AI aims to provide a more polished and streamlined voice input experience by intelligently editing out common speech impediments like "ums" and "uhs," as well as self-corrections. The technology is set to be integrated across the broader Google ecosystem, building on its initial debut in features like the Gboard "Rambler" on Pixel 11 phones.
According to Google, Gemini 3.5 Transcribe represents a substantial leap forward compared to its predecessor, Chirp 3. The company reports that the new model is approximately 70 percent faster in converting spoken words to final transcribed text. Furthermore, the live-speech error rate has been reduced to 5.5 percent, a notable improvement over Chirp 3's 7.32 percent. While the reduction in errors might seem marginal, the ability to reduce the friction of correcting typos during voice input offers a significant user benefit.
Beyond mere word recognition, Gemini 3.5 Transcribe is engineered to better understand the user's intended meaning. It actively removes filler words and corrects verbal stumbles in real-time. A key feature is its ability to reference a user-provided custom vocabulary, enabling it to accurately transcribe specialized jargon. The model supports 85 languages and can process audio with up to three speakers in pre-recorded content. However, users should be aware that the AI does make subtle wording changes to achieve its polished output, which may not be suitable for all applications where verbatim accuracy is critical.
Expanded Availability and Developer Access
Gemini 3.5 Transcribe is already making its way to users and developers. The "Rambler" feature in Gboard, powered by this new AI, is currently exclusive to Pixel 11 devices but is slated for expansion to more Gemini Intelligence devices later this year. Users of the Gemini app on macOS will also benefit from the enhanced voice input starting immediately. Google is also prioritizing developer access, with the model integrated into Antigravity starting today, offering full access to screen context and chat history with user permission. AI Studio's build model now includes Gemini 3.5 Transcribe, empowering developers to create AI-optimized voice-to-text applications. Furthermore, developers can access the model via the Gemini API. For the general public not using these specific platforms, Google plans to roll out the AI transcription feature to the Chrome browser soon. This integration will enable voice input for text fields across the web, facilitating tasks like composing emails, crafting comments, or interacting with AI chatbots.
The introduction of Gemini 3.5 Transcribe underscores Google's ongoing commitment to advancing AI-driven communication tools. By refining the nuances of human speech, Google aims to make voice interactions more natural, efficient, and accessible. This development is particularly relevant in an era where hands-free operation and voice commands are becoming increasingly integrated into daily digital life. The ability to reduce transcription errors and speed up the process could significantly impact productivity for professionals, students, and casual users alike, especially as voice interfaces become more sophisticated and widely adopted.
