English

WanSong v1.0 Technical Report

Audio and Speech Processing 2026-07-16 v1 Computer Vision and Pattern Recognition

Abstract

Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present \textbf{WanSong}, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), \textbf{WanSong} is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.

Cite

@article{arxiv.2607.14749,
  title  = {WanSong v1.0 Technical Report},
  author = {Binghui Chen and Pandeng Li and Yu Liu and Jingren Zhou},
  journal= {arXiv preprint arXiv:2607.14749},
  year   = {2026}
}

Comments

Wan Team