English

A2SB: Audio-to-Audio Schrodinger Bridges

Sound 2025-08-14 v2 Machine Learning Audio and Speech Processing

Abstract

Real-world audio is often degraded by numerous factors. This work presents an audio restoration model tailored for high-res music at 44.1kHz. Our model, Audio-to-Audio Schr\"odinger Bridges (A2SB), is capable of both bandwidth extension (predicting high-frequency components) and inpainting (re-generating missing segments). Critically, A2SB is end-to-end requiring no vocoder to predict waveform outputs, able to restore hour-long audio inputs, and trained on permissively licensed music data. A2SB is capable of achieving state-of-the-art band-width extension and inpainting quality on several out-of-distribution music test sets.

Keywords

Cite

@article{arxiv.2501.11311,
  title  = {A2SB: Audio-to-Audio Schrodinger Bridges},
  author = {Zhifeng Kong and Kevin J Shih and Weili Nie and Arash Vahdat and Sang-gil Lee and Joao Felipe Santos and Ante Jukic and Rafael Valle and Bryan Catanzaro},
  journal= {arXiv preprint arXiv:2501.11311},
  year   = {2025}
}
R2 v1 2026-06-28T21:11:03.602Z