English

ALICE: A Multifaceted Evaluation Framework of Large Audio-Language Models' In-Context Learning Ability

Sound 2026-03-24 v1 Artificial Intelligence Computation and Language Audio and Speech Processing

Abstract

While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains unstudied. To address this gap, we present ALICE, a three-stage framework that progressively reduces textual guidance to systematically evaluate LALMs' in-context learning ability under audio conditioning. Evaluating six LALMs across four audio understanding tasks under two output constraint categories, we uncover a consistent asymmetry across all stages and LALMs: in-context demonstrations reliably improve format compliance but fail to improve, and often degrade, the core task performance. This suggests that LALMs can glean surface-level formatting patterns from demonstrations but may struggle to leverage cross-modal semantic grounding to reliably infer task objectives from audio-conditioned examples, highlighting potential limitations in current cross-modal integration.

Keywords

Cite

@article{arxiv.2603.20433,
  title  = {ALICE: A Multifaceted Evaluation Framework of Large Audio-Language Models' In-Context Learning Ability},
  author = {Yen-Ting Piao and Jay Chiehen Liao and Wei-Tang Chien and Toshiki Ogimoto and Shang-Tse Chen and Yun-Nung Chen and Chun-Yi Lee and Shao-Yuan Lo},
  journal= {arXiv preprint arXiv:2603.20433},
  year   = {2026}
}

Comments

Submitted to Interspeech 2026

R2 v1 2026-07-01T11:30:37.773Z