BytePlus Voice Cloning

Upload a 10-15 second sample of a person's voice and the model learns the timbre, then reads any text you provide in that cloned voice.

Upload a 10-15 second sample of a person's voice and the model learns the timbre, then reads any text you provide in that cloned voice. Cloning defaults to Chinese, so a clean Chinese voice sample gives the best result. Use it for brand spokesperson audio, recurring narration, or creating consistent voice-overs. Runs through BytePlus; follow applicable consent and copyright rules.