English

Amphion: An Open-Source Audio, Music and Speech Generation Toolkit

Sound 2024-09-17 v3 Audio and Speech Processing

Abstract

Amphion is an open-source toolkit for Audio, Music, and Speech Generation, targeting to ease the way for junior researchers and engineers into these fields. It presents a unified framework that includes diverse generation tasks and models, with the added bonus of being easily extendable for new incorporation. The toolkit is designed with beginner-friendly workflows and pre-trained models, allowing both beginners and seasoned researchers to kick-start their projects with relative ease. The initial release of Amphion v0.1 supports a range of tasks including Text to Speech (TTS), Text to Audio (TTA), and Singing Voice Conversion (SVC), supplemented by essential components like data preprocessing, state-of-the-art vocoders, and evaluation metrics. This paper presents a high-level overview of Amphion. Amphion is open-sourced at https://github.com/open-mmlab/Amphion.

Keywords

Cite

@article{arxiv.2312.09911,
  title  = {Amphion: An Open-Source Audio, Music and Speech Generation Toolkit},
  author = {Xueyao Zhang and Liumeng Xue and Yicheng Gu and Yuancheng Wang and Jiaqi Li and Haorui He and Chaoren Wang and Songting Liu and Xi Chen and Junan Zhang and Zihao Fang and Haopeng Chen and Tze Ying Tang and Lexiao Zou and Mingxuan Wang and Jun Han and Kai Chen and Haizhou Li and Zhizheng Wu},
  journal= {arXiv preprint arXiv:2312.09911},
  year   = {2024}
}

Comments

Accepted by IEEE SLT 2024

R2 v1 2026-06-28T13:52:35.280Z