🔊Google DeepMind Rolls Out Gemini 3.1 Flash TTS With Directable Audio Tags
Google DeepMind Rolls Out Gemini 3.1 Flash TTS With Directa…
TL;DR
Google DeepMind launched Gemini 3.1 Flash TTS, its most controllable text-to-speech model yet.
Google DeepMind launched Gemini 3.1 Flash TTS, its most controllable text-to-speech model yet. New Audio Tags let developers direct vocal style, delivery, and pace through in-prompt text commands. The model is available via the Gemini API, Google AI Studio, Vertex AI for enterprises, and Google Vids for consumers.

Key Points
Audio Tags offer granular control over vocal style and delivery
Accessible via Gemini API, AI Studio, Vertex AI, and Google Vids
Framed as Google's most controllable TTS model to date
Why It Matters
Programmable voice control changes the economics of voice AI, audiobook production, and localization at scale, tightening the competitive race with ElevenLabs.
Comments
Be the first to comment
Enjoyed this article?
Get it daily. 7am. Free. Reads in 5 minutes.
Join 3,484 builders reading daily.