English

MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses

Audio and Speech Processing 2025-07-09 v1

Abstract

Smart glasses are increasingly positioned as the next-generation interface for ubiquitous access to large language models (LLMs). Nevertheless, achieving reliable interaction in real-world noisy environments remains a major challenge, particularly due to interference from side speech. In this work, we introduce a novel side-talk rejection multi-microphone Whisper (MMW) framework for smart glasses, incorporating three key innovations. First, we propose a Mix Block based on a Tri-Mamba architecture to effectively fuse multi-channel audio at the raw waveform level, while maintaining compatibility with streaming processing. Second, we design a Frame Diarization Mamba Layer to enhance frame-level side-talk suppression, facilitating more efficient fine-tuning of Whisper models. Third, we employ a Multi-Scale Group Relative Policy Optimization (GRPO) strategy to jointly optimize frame-level and utterance-level side speech suppression. Experimental evaluations demonstrate that the proposed MMW system can reduce the word error rate (WER) by 4.95\% in noisy conditions.

Cite

@article{arxiv.2507.05609,
  title  = {MMW: Side Talk Rejection Multi-Microphone Whisper on Smart Glasses},
  author = {Yang Liu and Li Wan and Yiteng Huang and Yong Xu and yangyang shi and Saurabh Adya and ming sun and Florian Metze},
  journal= {arXiv preprint arXiv:2507.05609},
  year   = {2025}
}
R2 v1 2026-07-01T03:50:40.632Z