Google's Gemini AI Transcribe Ditches 'Ums' and 'Ahs'
Google's Gemini Audio introduces Gemini 3.5 Transcribe, an AI that automatically removes filler words and recognizes specialized jargon in over 85 languages. The tool aims for more natural voice editing.

Google has unveiled a significant upgrade to its Gemini Audio platform with the introduction of Gemini 3.5 Transcribe. This new artificial intelligence model boasts enhanced capabilities, including the automatic detection of specialized jargon and support for over 85 languages. Gemini 3.5 Transcribe represents a substantial leap forward from its predecessor, Chirp 3, with notable improvements in multilingual performance and reduced wording errors, according to Google.
A key feature of Gemini 3.5 Transcribe is its ability to allow users to "edit naturally with just your voice," as stated by the company. The AI can automatically format transcribed text and intelligently remove common filler words such as "um" and "uh." For professionals and users dealing with specific terminology, the model offers a customization option. Users can provide a tailored vocabulary list, enabling Gemini 3.5 Transcribe to adapt to unique spellings and industry-specific jargon, thereby preventing the need for manual corrections.
Beyond text refinement, the AI can also attribute speech to up to three distinct speakers within pre-recorded audio files. This is complemented by precise word-level timestamps, offering granular control over the transcribed content. This advanced AI functionality is designed to streamline audio editing and transcription workflows for a variety of applications, from content creation to meeting summaries.
Context and Broader Implications
The launch of Gemini 3.5 Transcribe arrives as Google continues to expand its AI offerings, with users still anticipating the full release of the Gemini 3.5 Pro model, originally slated for June. This development underscores a broader trend in the tech industry towards more sophisticated and user-friendly AI tools that integrate seamlessly into daily workflows. The ability to process audio with such accuracy and context-awareness has far-reaching implications for accessibility, productivity, and communication across different sectors.
Google also mentioned updates to Gemini 3.5 Live and Gemini 3.5 Live Experimental, though it later clarified that only Gemini 3.5 Transcribe is being announced today, with no new launch date provided for the other models. Previously, Gemini 3.5 Live was described as being more adept at handling interruptions mid-sentence, language recognition, and live visual processing. Gemini 3.5 Live Experimental was intended to provide real-time narration of its progress on complex reasoning tasks.
Gemini 3.5 Transcribe is commencing its rollout today, initially available in English for all macOS Gemini app users. Additionally, the Rambler dictation feature on Android will receive this update in select countries and languages. Developers can access the model through a public preview in the Gemini API via AI Studio and Antigravity. Google has indicated that support for Chrome is planned for the near future, further expanding its accessibility.
The advancements in transcription technology, particularly in recognizing diverse languages and specialized vocabulary, highlight the increasing power of AI in breaking down communication barriers. Tools like Gemini 3.5 Transcribe have the potential to democratize access to information and enhance collaboration by providing accurate and efficient transcription services for a global audience.
The ability to customize the vocabulary is a significant step towards personalized AI experiences. As AI models become more integrated into everyday tools, features that cater to niche terminology and specific user needs will become increasingly important for widespread adoption and effectiveness.
