English

Seed-Music: A Unified Framework for High Quality and Controlled Music Generation

Sound 2024-09-20 v3 Audio and Speech Processing

Abstract

We introduce Seed-Music, a suite of music generation systems capable of producing high-quality music with fine-grained style control. Our unified framework leverages both auto-regressive language modeling and diffusion approaches to support two key music creation workflows: controlled music generation and post-production editing. For controlled music generation, our system enables vocal music generation with performance controls from multi-modal inputs, including style descriptions, audio references, musical scores, and voice prompts. For post-production editing, it offers interactive tools for editing lyrics and vocal melodies directly in the generated audio. We encourage readers to listen to demo audio examples at https://team.doubao.com/seed-music "https://team.doubao.com/seed-music".

Keywords

Cite

@article{arxiv.2409.09214,
  title  = {Seed-Music: A Unified Framework for High Quality and Controlled Music Generation},
  author = {Ye Bai and Haonan Chen and Jitong Chen and Zhuo Chen and Yi Deng and Xiaohong Dong and Lamtharn Hantrakul and Weituo Hao and Qingqing Huang and Zhongyi Huang and Dongya Jia and Feihu La and Duc Le and Bochen Li and Chumin Li and Hui Li and Xingxing Li and Shouda Liu and Wei-Tsung Lu and Yiqing Lu and Andrew Shaw and Janne Spijkervet and Yakun Sun and Bo Wang and Ju-Chiang Wang and Yuping Wang and Yuxuan Wang and Ling Xu and Yifeng Yang and Chao Yao and Shuo Zhang and Yang Zhang and Yilin Zhang and Hang Zhao and Ziyi Zhao and Dejian Zhong and Shicen Zhou and Pei Zou},
  journal= {arXiv preprint arXiv:2409.09214},
  year   = {2024}
}

Comments

Seed-Music technical report, 20 pages, 5 figures