English

Attention-guided Evidence Grounding for Spoken Question Answering

Computation and Language 2026-03-19 v2 Artificial Intelligence

Abstract

Spoken Question Answering (Spoken QA) presents a challenging cross-modal problem: effectively aligning acoustic queries with textual knowledge while avoiding the latency and error propagation inherent in cascaded ASR-based systems. In this paper, we introduce Attention-guided Evidence Grounding (AEG), a novel end-to-end framework that leverages the internal cross-modal attention of Speech Large Language Models (SpeechLLMs) to explicitly locate and ground key evidence in the model's latent space. To address the diffuse attention distribution in pre-trained models, we propose Learning to Focus on Evidence (LFE), a supervised fine-tuning paradigm that calibrates the model's attention mechanism to distinguish query-relevant segments from irrelevant context. Experiments on SQuAD, HotpotQA, and MuSiQue demonstrate that AEG reduces hallucinations and achieves strong efficiency gains, outperforming large-scale cascaded baselines (Whisper-Large-v3 + Reranker) while reducing inference latency by approximately 62%.

Keywords

Cite

@article{arxiv.2603.16292,
  title  = {Attention-guided Evidence Grounding for Spoken Question Answering},
  author = {Ke Yang and Bolin Chen and Yuejie Li and Yueying Hua and Jianhao Nie and Yueping He and Bowen Li and Chengjun Mao},
  journal= {arXiv preprint arXiv:2603.16292},
  year   = {2026}
}

Comments

Accepted to ICME 2026