Google Gemini 2.5 Flash Text to Speech

Quickly turn text into natural speech with Google Gemini 2.5 Flash TTS; choose from 30 voices and set the spoken language.

Google Gemini 2.5 Flash Text to Speech is Google's low-latency, cost-effective speech model that quickly turns your text into natural, fluent spoken audio, well suited to high-volume or near-real-time voice-over needs. Pick one of 30 prebuilt voices from the dropdown, each with its own character such as bright, firm, warm, or informative, and choose the spoken language from the second dropdown (auto-detect by default; Chinese, English, Japanese, Korean, French, German, Spanish and many more are supported). You can also steer the delivery in plain language, for example by starting with "Say calmly:" or "Say cheerfully:", and the model adjusts emotion, pace and tone accordingly. When you need the highest fidelity and the most refined prosody, switch to Gemini 2.5 Pro Text to Speech. Great for short-video narration, customer-service prompts, course reading, announcements and accessibility audio. Text input is limited to 8,192 tokens and the audio is delivered as MP3. Maximum context window (input and output combined): about 24,576 tokens. Maximum total input per request (including history): 8,192 tokens. Maximum output per reply: about 16,384 tokens (model reasoning, if any, counts toward this limit). Output content formats: Audio Audio output format: MP3. This module does not support continuous conversations. AI can make mistakes. Please verify important information.