According to Beating, ByteDance's Seed Audio 1.0 was launched on July 22, enabling single-prompt generation of dialogue, sound effects, and ambient audio with 100ms timing precision. The model supports text and reference audio inputs, can produce approximately two minutes of audio per generation with extended capability, and maintains consistent character voices across outputs.
The audio demonstrates over 90% usability in most scenarios, with Mean Opinion Score (MOS) exceeding 4 across most supported languages. Seed Audio 1.0 supports 20+ languages with cross-lingual voice transfer and tone customization. The model is now available through Volcano Ark experience center with API integration open.