BytePlus Voice Cloning

Upload a 10-15 second sample of a person's voice and the model learns the timbre, then reads any text you provide in that cloned voice.

Upload a 10-15 second sample of a person's voice and the model learns the timbre, then reads any text you provide in that cloned voice. Cloning defaults to Chinese, so a clean Chinese voice sample gives the best result. Use it for brand spokesperson audio, recurring narration, or creating consistent voice-overs. Runs through BytePlus; follow applicable consent and copyright rules. Maximum total upload size per request: 256MB (all files and message content combined). Output content formats: Audio Audio output format: MP3. This module does not support continuous conversations. AI can make mistakes. Please verify important information.