English

Fast and Flexible Audio Bandwidth Extension via Vocos

Audio and Speech Processing 2026-03-10 v1 Machine Learning Sound

Abstract

We propose a Vocos-based bandwidth extension model that enhances audio at 8-48 kHz by generating missing high-frequency content. Inputs are resampled to 48 kHz and processed by a neural vocoder backbone, enabling a single network to support arbitrary upsampling ratios. A lightweight Linkwitz-Riley-inspired refiner merges the original low band with the generated high frequencies via a smooth crossover. On validation, the model achieves competitive log-spectral distance while running at a real-time factor of 0.0001 on an NVIDIA A100 GPU and 0.0053 on an 8-core CPU, demonstrating practical, high-quality BWE at extreme throughput.

Keywords

Cite

@article{arxiv.2603.07285,
  title  = {Fast and Flexible Audio Bandwidth Extension via Vocos},
  author = {Yatharth Sharma},
  journal= {arXiv preprint arXiv:2603.07285},
  year   = {2026}
}

Comments

5 pages, 2 figures, 5 tables. Submitted to INTERSPEECH 2026. Code available at https://github.com/ysharma3501/LavaSR.git