Google launches Gemini 3.5 Transcribe, its fastest speech to text model supporting 85 languages
Listen to this article
Read by Anchor
Seventy percent. That is the improvement Gemini 3.5 Transcribe recorded in time to final transcript compared with the previous generation, according to independent testing by Artificial Analysis. A single figure captures a significant shift in speech to text, positioning the new model as a serious contender for any developer building voice applications today.
What changed substantially
Google announced Gemini 3.5 Transcribe on its official blog, describing it as its most accurate speech to text model to date, designed specifically for intelligent voice interactions. The model goes beyond accuracy, adding three capabilities that until recently required separate models or complex engineering: developer defined custom vocabulary recognition, automatic support for more than 85 languages with regional dialects, and automated filler word removal and text formatting.
For developers: two interfaces, not one
Immediate availability comes through two routes: the Gemini app on macOS for general users, and the Gemini API in Google AI Studio and Antigravity for developers, both in public preview. The second route is technically the more significant: it allows integrating live streaming and prerecorded transcription into applications without managing separate infrastructure. Companies such as Fevo, Intellect Health, and LingoPal told Google that they observed low latency, high accuracy, and broad language support in early testing.
Accuracy figures reported by Google
The official blog points to a word error rate (WER) of 4.0% in live streaming mode and 2.6% in recorded file mode, alongside support for identifying up to three speakers with timestamps. These figures, if they hold up in real production environments, put the model in direct competition with specialized solutions companies have relied on for years. The 70% improvement in time to final transcript, as measured by Artificial Analysis, means interactive voice applications, from personal assistants to automated contact centers, can now respond with a fluidity previously available only with heavier, more expensive models.
Integration points for regional developers
The public preview via the Gemini API allows immediate integration without waiting for approvals or waitlists. Product teams in Riyadh, Dubai, and Cairo can now test the model with their own vocabularies, product names, medical terms, banking codes, and see how accurately it recognizes them in a mixed Arabic environment. Google AI Studio documentation provides ready code samples for live streaming and recorded files, reducing prototyping time from weeks to hours.
The bet on Arabic and regional dialects
The angle that matters to readers in the Gulf, Egypt, and the Levant is not just the 85 languages figure. The real value lies in the seamless handling of diverse regional dialects cited in the official announcement. Regional markets operate daily in Modern Standard Arabic alongside Gulf, Egyptian, and Levantine dialects, as well as English, French, and Urdu within the same workplace. A single model that understands this mix without switching settings saves months of work for product teams that currently have to assemble specialized models for each dialect.
What remains unresolved
The official disclosure did not specify pricing for Gemini 3.5 Transcribe or hardware requirements to run it, nor did it list the 85 languages to explicitly confirm Arabic support, and it did not set a timeline for Chrome integration. These are genuine gaps rather than oversights, and any product team should account for them in its evaluation plan before committing.
The bottom line: a 70% improvement in speed, broad language support, and competitive accuracy figures place Gemini 3.5 Transcribe among the tools worth testing seriously this week rather than next year. For anyone building a voice product in Arabic or regional languages, the public preview is available now in Google AI Studio, leaving a single step between an idea and a working prototype.