English

Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds

Audio and Speech Processing 2024-09-30 v1 Artificial Intelligence Sound Signal Processing

Abstract

This paper introduces a speech enhancement solution tailored for true wireless stereo (TWS) earbuds on-device usage. The solution was specifically designed to support conversations in noisy environments, with active noise cancellation (ANC) activated. The primary challenges for speech enhancement models in this context arise from computational complexity that limits on-device usage and latency that must be less than 3 ms to preserve a live conversation. To address these issues, we evaluated several crucial design elements, including the network architecture and domain, design of loss functions, pruning method, and hardware-specific optimization. Consequently, we demonstrated substantial improvements in speech enhancement quality compared with that in baseline models, while simultaneously reducing the computational complexity and algorithmic latency.

Keywords

Cite

@article{arxiv.2409.18705,
  title  = {Speech Boosting: Low-Latency Live Speech Enhancement for TWS Earbuds},
  author = {Hanbin Bae and Pavel Andreev and Azat Saginbaev and Nicholas Babaev and Won-Jun Lee and Hosang Sung and Hoon-Young Cho},
  journal= {arXiv preprint arXiv:2409.18705},
  year   = {2024}
}

Comments

Accepted by Interspeech 2024

R2 v1 2026-06-28T18:59:28.097Z