English

Smule Renaissance Small: Efficient General-Purpose Vocal Restoration

Sound 2025-10-27 v1

Abstract

Vocal recordings on consumer devices commonly suffer from multiple concurrent degradations: noise, reverberation, band-limiting, and clipping. We present Smule Renaissance Small (SRS), a compact single-stage model that performs end-to-end vocal restoration directly in the complex STFT domain. By incorporating phase-aware losses, SRS enables large analysis windows for improved frequency resolution while achieving 10.5x real-time inference on iPhone 12 CPU at 48 kHz. On the DNS 5 Challenge blind set, despite no speech training, SRS outperforms a strong GAN baseline and closely matches a computationally expensive flow-matching system. To enable evaluation under realistic multi-degradation scenarios, we introduce the Extreme Degradation Bench (EDB): 87 singing and speech recordings captured under severe acoustic conditions. On EDB, SRS surpasses all open-source baselines on singing and matches commercial systems, while remaining competitive on speech despite no speech-specific training. We release both SRS and EDB under the MIT License.

Keywords

Cite

@article{arxiv.2510.21659,
  title  = {Smule Renaissance Small: Efficient General-Purpose Vocal Restoration},
  author = {Yongyi Zang and Chris Manchester and David Young and Ivan Ivanov and Jeffrey Lufkin and Martin Vladimirov and PJ Solomon and Svetoslav Kepchelev and Fei Yueh Chen and Dongting Cai and Teodor Naydenov and Randal Leistikow},
  journal= {arXiv preprint arXiv:2510.21659},
  year   = {2025}
}

Comments

Technical Report

R2 v1 2026-07-01T07:04:18.483Z