English

EvA: An Evidence-First Audio Understanding Paradigm for LALMs

Sound 2026-05-29 v2 Artificial Intelligence

Abstract

Large Audio Language Models (LALMs) still struggle in complex acoustic scenes because they often fail to preserve task-relevant acoustic evidence before reasoning begins. We identify this error pattern as the evidence bottleneck: state-of-the-art systems show larger deficits in acoustic evidence extraction than in downstream reasoning, suggesting that upstream perception is often the limiting factor. To address this problem, we propose EvA (Evidence-First Audio), a dual-path architecture that enhances acoustic evidence preservation through hierarchical aggregation and non-compressive, time-aligned fusion. We also build EvA-Perception, a large-scale training set with about 54K event-ordered captions and 500K evidence-grounded QA pairs. Under a unified zero-shot protocol, EvA achieves the best open-source \emph{Perception} results on MMAU, MMAR, and MMSU, with the largest gains on perception-heavy splits. Human evaluation on open-ended captioning further shows improved fine-grained acoustic coverage and caption quality. These results support the evidence-first hypothesis: stronger audio understanding depends on preserving acoustic evidence before reasoning. Project can be found at https://satsuki2486441738.github.io/EvA/.

Keywords

Cite

@article{arxiv.2603.27667,
  title  = {EvA: An Evidence-First Audio Understanding Paradigm for LALMs},
  author = {Xinyuan Xie and Shunian Chen and Zhiheng Liu and Yuhao Zhang and Zhiqiang Lv and Liyin Liang and Benyou Wang},
  journal= {arXiv preprint arXiv:2603.27667},
  year   = {2026}
}
R2 v1 2026-07-01T11:42:51.969Z