Google has launched its most accurate speech-to-text model to date, moving from "word-for-word transcription" to "understanding-based transcription."
On August 26 local time, Google released its latest speech-to-text model, Gemini 3.5 Transcribe. This model not only boasts higher recognition accuracy, but also extends its capabilities from word-for-word recording to understanding the speaker's intent and automatically organizing the output results , marking a step forward for speech recognition technology towards deeper language understanding.
Unlike traditional speech recognition tools, Gemini 3.5 Transcribe can recognize the speaker's self-correction, automatically remove filler words, and directly convert unstructured spoken language into formatted text.
Google calls it the "most accurate" speech-to-text model to date and has already introduced it into the Android version of Gboard and the Mac version of Gemini, while also making it available for preview to developers through the Gemini API.
This announcement, released on the same day as the Gemini 3.5 Live and Gemini 3.5 Live Experimental, together form Google's new Gemini Audio series.
Analysts point out that as voice input rapidly penetrates high-frequency products such as Google's Chrome and Gboard, voice interaction is expected to gradually replace keyboard input and become the main way for users to connect with AI systems.
As of press time, Google's stock price was down 1.37% during Thursday's trading session.

From "verbal transcription" to "intent understanding"
The core upgrade of Gemini 3.5 Transcribe lies in its ability to understand the context of natural language. Analysis indicates that the key to this release is not just improving traditional recognition accuracy, but enabling the model to understand the speaker's intent during transcription and "edit" the original speech.
Specifically, when a user says "Tuesday—no, Wednesday," the model can recognize that this is a self-correction rather than recording both statements.
At the same time, the model automatically removes filler words like "um" and "uh" and tidies up punctuation and text formatting. This model can directly convert messy, unstructured spoken content into formatted text.
In terms of language coverage, the model supports over 85 languages, features automatic language recognition, and can handle multilingual scenarios. Google also allows users to add custom vocabulary, enabling the model to more accurately recognize technical terms, special spellings, and alphanumeric combinations such as order numbers and zip codes. For pre-recorded audio, the model also supports speaker recognition and word-by-word timestamps.
Simultaneous advancement of developer APIs and product implementation
In terms of open capabilities to developers, Gemini 3.5 Transcribe has entered the preview stage through the Gemini API, supporting two processing scenarios: real-time voice streaming and recorded audio files.
According to Google's official documentation, document transcription can be up to one hour long, while real-time transcription is designed for low-latency applications. Developers can utilize features such as low-latency transcription, custom vocabulary, and intelligent formatting to embed speech recognition capabilities into their own applications.
On the product side, Gemini 3.5 Transcribe has been implemented on the Android platform with the Gboard Rambler voice input function and on the Mac version of the Gemini application.
Google also plans to bring this to the Chrome browser, at which point users will be able to directly input content via voice in web page text input boxes, covering scenarios such as replying to messages, writing posts, and issuing commands to Gemini.
The Business Logic Behind the Battle for Voice Access
This release is part of Google's systematic advancement of its Gemini "voice gateway" strategy.
Google is extending Gemini's capabilities beyond text and images to include real-time conversations, voice translation, and voice input. The Gemini 3.5 Transcribe, Gemini 3.5 Live, and Gemini 3.5 Live Experimental together form its Gemini Audio series.
From a commercial perspective, speech-to-text technology is evolving from a simple backend capability into a crucial interactive entry point for AI applications. As models acquire integrated capabilities of "understanding, comprehending, organizing, and executing," the way users interact with AI is expected to shift from keyboard input to voice commands.
For Google, embedding Gemini 3.5 Transcribe into frequently used products such as Chrome, Gboard, Docs, and Gmail is expected to further expand Gemini's penetration in daily productivity scenarios and consolidate its competitive position in the emerging field of voice interaction.
Risk Warning and DisclaimerInvesting involves risk; please exercise caution. This article does not constitute personal investment advice and does not take into account the specific investment objectives, financial situation, or needs of individual users. Users should consider whether any opinions, views, or conclusions in this article are suitable for their specific circumstances. Any investment decisions made based on this information are at your own risk.