English

Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models

Audio and Speech Processing 2026-04-21 v2

Abstract

Recent advances in reasoning models have driven significant progress in text and multimodal domains, yet audio reasoning remains relatively limited. Only a few Large Audio Language Models (LALMs) incorporate explicit Chain-of-Thought (CoT) reasoning, and their capabilities are often inconsistent and insufficient for complex tasks. To bridge this gap, we introduce Audio-Cogito, a fully open-source solution for deep audio reasoning. We develop Cogito-pipe for high-quality audio reasoning data curation, producing 545k reasoning samples that will be released after review. Based on this dataset, we adopt a self-distillation strategy for model fine-tuning. Experiments on the MMAR benchmark, the only audio benchmark evaluating the CoT process, show that our model achieves the best performance among open-source models and matches or surpasses certain closed-source models in specific metrics. Our approach also ranks among the top-tier systems in the Interspeech 2026 Audio Reasoning Challenge.

Keywords

Cite

@article{arxiv.2604.12527,
  title  = {Audio-Cogito: Towards Deep Audio Reasoning in Large Audio Language Models},
  author = {Longhao Li and Hongjie Chen and Zehan Li and Qihan Hu and Jian Kang and Jie Li and Lei Xie and Yongxiang Li},
  journal= {arXiv preprint arXiv:2604.12527},
  year   = {2026}
}

Comments

Submitted to Interspeech 2026

R2 v1 2026-07-01T12:08:26.871Z