Google's Gemini 3.1 Flash TTS speaks 70+ languages with audio tags
Google's new TTS model hits 70+ languages, natural-language voice controls, and SynthID watermarking — available now in AI Studio and the Gemini API.

Google just shipped Gemini 3.1 Flash TTS, its most capable text-to-speech model yet. It supports 70+ languages with fine-grained control over accent, pacing, and style — all steerable through plain English inline audio tags embedded directly in your text input.
The model scores 1,211 on the Artificial Analysis TTS leaderboard and sits in the quality-to-price sweet spot for developers. Multi-speaker dialogue is native, and you can export voice settings as API code for consistent reproduction across projects. All output is invisibly watermarked with SynthID to flag AI-generated audio.
It's live today in Google AI Studio, the Gemini API (developer preview), Vertex AI (enterprise preview), and Google Vids for Workspace users.
Why it matters: Natural-language voice control + 70-language coverage + anti-deepfake watermarking in one affordable API call is a serious leap for anyone building voice-first AI products.
Sources
Independent coverage
- Gemini 3.1 Flash TTS: New text-to-speech AI model blog.google
- Google's Gemini 3.1 Flash TTS offers unparalleled control over AI voices siliconangle.com
Written by an AI pipeline from the sources above. Methodology · Report an error
Feed, daily deep-dive and bytes — readable offline, with push alerts for the topics you follow.