Google on Wednesday unveiled Gemini 3.5 Transcribe, a new AI model designed for precise, real-time speech-to-text conversion that goes far beyond standard voice recognition by understanding context, formatting, and user intent.
The model, announced on Google’s Keyword blog, converts raw audio directly into polished, formatted text without the preprocessing steps that traditional speech recognition systems require. It handles background noise, complex jargon, and filler words like “ums” and “ahs” in a single pass, producing clean output at the speed of speech.
Gemini 3.5 Transcribe is already live across several Google products. On Android, it powers a new feature called Rambler within the Gboard keyboard, which transforms spoken thoughts into well-formatted text while filtering out disfluencies. The model also runs inside the Gemini app on both Android and macOS, and Google confirmed it is coming to Chrome for post-call analytics and dictation workflows.
Developer Access and Benchmarks
For developers, the model is available through two separate APIs: the Gemini API in Google AI Studio for streaming and batch transcription, and the Gemini Enterprise Agent Platform for production deployments. Google said the model supports more than 85 languages and is designed to plug directly into existing developer workflows for voice agents, real-time captioning, and call center analytics.
On the FLEURS benchmark, a standard evaluation for multilingual speech recognition, Gemini 3.5 Transcribe achieved a 5.50 percent word error rate in streaming mode and 5.04 percent in non-streaming use cases. Google said these results represent an improvement over Chirp 3, its previous speech recognition model, across a set of top languages and locales.
The model also introduces function calling capabilities, which allow it to delegate complex tasks to other Gemini models. Users can trigger image generation, live search, summarization, and file analysis directly from voice input, creating a unified voice interface across Google’s AI ecosystem.
Competing in a Crowded Market
The launch positions Google against a growing field of companies investing in speech-to-text AI. OpenAI, Microsoft, and Amazon all offer competing transcription services, but Google’s integration with its consumer products gives it an immediate scale advantage. By embedding Gemini 3.5 Transcribe into Gboard, which ships on billions of Android devices, Google can drive adoption through everyday use rather than relying solely on enterprise contracts.
Speech is becoming an increasingly important interface for AI systems. Voice agents, call-center automation, meeting transcription tools, accessibility software, and hands-free computing all depend on accurate, low-latency speech recognition. The market for voice AI technology is projected to grow significantly as more applications move beyond text-based chat interfaces.
Google’s approach differs from traditional speech recognition pipelines in several ways. Conventional systems break audio into small segments, transcribe each independently, then attempt to assemble a coherent output. Gemini 3.5 Transcribe processes the entire audio stream at once, using its multimodal capabilities to maintain context across the full recording. This eliminates errors that typically occur at segment boundaries and allows the model to apply formatting rules, such as punctuation and paragraph breaks, consistently throughout.
Adaptable Editing and Custom Vocabulary
The model also supports what Google calls adaptable voice editing, allowing users to refine their spoken output in real time. A speaker can correct details, clarify spellings, or change writing style without restarting the transcription process. For specialized industries, the system can be taught custom vocabulary to recognize brand names, product titles, and technical terminology, reducing the misspellings and misunderstandings that plague generic transcription tools.
“The transcription model is not only fast, but impressively accurate with emails and numbers, two areas where the industry still needs meaningful advancement,” Google said in its announcement.
The release comes as Google prepares for its next major Gemini model update. The company has not yet confirmed when Gemini 3.5 Pro will launch, but the introduction of Transcribe suggests Google is expanding its Gemini family with specialized, task-focused models rather than relying solely on general-purpose language models.
For Google, the strategic advantage lies in distribution. Gboard reaches billions of users worldwide, and the Chrome browser commands more than 60 percent of the global browser market. Embedding high-quality transcription into these surfaces creates a feedback loop: more usage generates more training data, which improves the model, which attracts more users and developers to the platform.
discussion