Google Gemini 2.5 Pro Text to Speech uses Google's highest-quality speech model to read your text aloud with clear, natural, expressive delivery. Pick one of 30 prebuilt voices from the dropdown, each with its own character such as bright, firm, warm, or informative, and choose the spoken language from the second dropdown (auto-detect by default; Chinese, English, Japanese, Korean, French, German, Spanish and many more are supported). You can also steer the delivery in plain language, for example by starting with "Say excitedly:" or "Whisper:", and the model adjusts emotion, pace and tone accordingly. Split very long content into parts, as quality may degrade on multi-minute continuous narration. Great for audiobooks, e-learning narration, podcast intros, product explainers and accessibility audio. Text input is limited to 8,192 tokens and the audio is delivered as MP3. Maximum context window (input and output combined): about 24,576 tokens. Maximum total input per request (including history): 8,192 tokens. Maximum output per reply: about 16,384 tokens (model reasoning, if any, counts toward this limit). Output content formats: Audio Audio output format: MP3. This module does not support continuous conversations. AI can make mistakes. Please verify important information.