Skip to content
SiliconANGLE·

🔊Google DeepMind Rolls Out Gemini 3.1 Flash TTS With Directable Audio Tags

Google DeepMind Rolls Out Gemini 3.1 Flash TTS With Directa…

TL;DR

Google DeepMind launched Gemini 3.1 Flash TTS, its most controllable text-to-speech model yet.

Google DeepMind launched Gemini 3.1 Flash TTS, its most controllable text-to-speech model yet. New Audio Tags let developers direct vocal style, delivery, and pace through in-prompt text commands. The model is available via the Gemini API, Google AI Studio, Vertex AI for enterprises, and Google Vids for consumers.

Google DeepMind Rolls Out Gemini 3.1 Flash TTS With Directable Audio Tags — SiliconANGLE

Key Points

1

Audio Tags offer granular control over vocal style and delivery

2

Accessible via Gemini API, AI Studio, Vertex AI, and Google Vids

3

Framed as Google's most controllable TTS model to date

Why It Matters

Programmable voice control changes the economics of voice AI, audiobook production, and localization at scale, tightening the competitive race with ElevenLabs.

GoogleGeminiTTSvoice AI

Frequently Asked Questions

Why does this matter?

Programmable voice control changes the economics of voice AI, audiobook production, and localization at scale, tightening the competitive race with ElevenLabs.

What happened?

Google DeepMind launched Gemini 3.1 Flash TTS, its most controllable text-to-speech model yet.

Comments

Subscribe to join the conversation...

Be the first to comment

Enjoyed this article?

Get it daily. 7am. Free. Reads in 5 minutes.

Join 3,461 builders reading daily.

Also get