BytePlus SeedAudio 1.0 generates audio from a natural-language description, covering ambient sound effects, narration, character lines, and atmospheric soundscapes. For example, type "Inside a football stadium, the crowd erupts as the commentator shouts: What a goal!" and you get a complete clip with crowd ambience and the spoken line. More than 20 languages are supported, including English, Chinese, Japanese, Korean, French, and German. Besides text-only generation, you can upload reference audio clips (up to 3 clips, each within 30 seconds and 10 MB, in wav, mp3, ogg, or opus format) and reference them in the prompt as @Audio1, @Audio2 in upload order so the generated voice follows the reference timbre; or upload one image instead (jpeg, png, or webp, up to 10 MB), in which case the text is read aloud in a style matching the picture. Audio and image references cannot be combined. The text prompt is limited to 3000 characters. Great for audiobooks, video dubbing, game sound effects, ad voice-overs, and short-form social clips. Generated audio is up to 120 seconds long and is delivered as MP3. Maximum total upload size per request: 256MB (all files and message content combined). Accepted image formats: jpeg, jpg, png, webp. Maximum number of images: 1. Maximum size per image: 10MB. Accepted audio formats: mp3, ogg, opus, wav. Maximum number of audio files: 3. Maximum size per audio file: 10MB. Output content formats: Audio Maximum text input: 3,000 characters. Audio output format: MP3 (up to 120 seconds). This module does not support continuous conversations.