Google introduces Gemini 3.5 Transcribe for live and batch speech transcription
Google has introduced the Gemini 3.5 Transcribe model for ongoing audio transcription and recording processing. The public preview is available through the Gemini API and select company products.

On August 26, Google introduced Gemini 3.5 Transcribe, a new speech-to-text model designed to transcribe speech in real time and from audio recordings. The company made it available in public preview through the Gemini API in Google AI Studio and through the Gemini Enterprise Agent Platform.
The new offering has two separate variants. The gemini-3.5-transcribe-live model is designed for ongoing transcription through the Gemini Live API. Its documentation cites low latency and streamed text output. The second model, gemini-3.5-transcribe, is intended for processing recordings and supports assigning individual parts of a recording to speakers, as well as word-level timestamps.
Gemini 3.5 Transcribe offers two transcription modes
With the new interface, Google combines speech recognition itself with options for subsequent text editing. The API supports automatic language detection, a custom vocabulary, and two output modes: verbatim and edited transcription.
The verbatim mode is intended to preserve spoken language, including filler words. The edited mode, also known as Smart transcription, can remove filler words and adjust the formatting of the resulting text. This output may be practical, for example, for meeting notes or dictation, but it may not suit situations where an exact verbatim record of a statement is required.
Slovak is also listed in the supported-language documentation under the sk-SK designation. However, Google does not provide separate accuracy results for Slovak. Practical quality may depend on the recording quality, background noise, specialized terms used, the number of speakers, and the selected configuration.
Live transcription and recording analysis
The Live variant targets applications that need text during an ongoing conversation or dictation. The streamed output could be used by a voice agent, for captioning, or in a user interface displaying an ongoing transcription.
The recording model, by contrast, is intended to process existing audio. In addition to text, it is supposed to provide word-level timestamps and assign segments to individual speakers. These features are relevant to creating transcripts of interviews, podcasts, or work meetings.
Google says Gemini 3.5 Transcribe is also available in the Gemini app for macOS, currently in English. The company is rolling out the Rambler feature in the Gboard keyboard for Android, but only in selected countries and languages. Integration into the Chrome browser is reportedly in preparation, with no confirmed date for general availability of dictation in web fields.
Performance claims have not yet been independently verified
In its announcement, Google refers to measurements by Artificial Analysis and attributes to them Word Error Rate figures between 2.6 and 4.0 percent, as well as a 70 percent acceleration in transcription finalization. These figures have not been independently reproduced in the available materials. It is likewise unconfirmed how the stated parameters will perform with Slovak speech or recordings involving a larger number of speakers.
For developers, the key point is that the interface is still in public preview. In the announced materials, Google does not provide complete information about pricing, usage limits, or the terms governing voice-data processing during this phase. These details, along with independent latency and accuracy tests, will be important when assessing deployment of the model in production services.
Sources
- Google Blog — Intelligent transcription with Gemini 3.5 Transcribe – Confirms the August 26, 2026 announcement, model variants, preview availability, product integrations, and properties stated by Google.
- Google AI for Developers — Live transcription with Gemini Live API – Confirms the `gemini-3.5-transcribe-live` model, streamed transcription, the technical interface, transcription features, and supported languages including Slovak.
- Google AI for Developers — Audio transcription – Documents the separate API for transcribing audio recordings.
Verified and updated: 08/27/2026 20:15



