English

Query-Based Asymmetric Modeling with Decoupled Input-Output Rates for Speech Restoration

Audio and Speech Processing 2026-01-29 v3

Abstract

Speech restoration in real-world conditions is challenging due to compounded distortions and mismatches between input and desired output rates. Most existing systems assume a fixed and shared input-output rate, relying on external resampling that incurs redundant computation and limits generality. We address this setting by formulating speech restoration under decoupled input-output rates, and propose TF-Restormer, a query-based asymmetric modeling framework. The encoder concentrates analysis on the observed input bandwidth using a time-frequency dual-path architecture, while a lightweight decoder reconstructs missing spectral content via frequency extension queries. This design enables a single model to operate consistently across arbitrary input-output rate pairs without redundant resampling. Experiments across diverse sampling rates, degradations, and operating modes show that TF-Restormer maintains stable restoration behavior and balanced perceptual quality, including in real-time streaming scenarios. Code and demos are available at https://tf-restormer.github.io/demo.

Keywords

Cite

@article{arxiv.2509.21003,
  title  = {Query-Based Asymmetric Modeling with Decoupled Input-Output Rates for Speech Restoration},
  author = {Ui-Hyeop Shin and Jaehyun Ko and Woocheol Jeong and Hyung-Min Park},
  journal= {arXiv preprint arXiv:2509.21003},
  year   = {2026}
}

Comments

Preprint. Under review

R2 v1 2026-07-01T05:55:51.177Z