English

Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking

Audio and Speech Processing 2025-11-04 v1

Abstract

In the era of large language models (LLMs) and artificial general intelligence (AGI), computer audition must evolve beyond traditional paradigms to fully leverage the capabilities of foundation models, towards more comprehensive understanding, more natural generation and more human-like interaction. Audio, as a modality rich in semantic, emotional, and contextual cues, plays a vital role in achieving naturalistic and embodied machine intelligence. This survey provides a comprehensive review of recent progress in integrating audio into LLMs, with a focus on four key areas: audio comprehension, audio generation, speech-based interaction, and audio-visual understanding. We analyze how LLMs are reshaping audio perception and reasoning, enabling systems to understand sound at a deeper semantic level, generate expressive audio outputs, and engage in human-like spoken interaction. Furthermore, we explore how the fusion of audio and visual modalities enhances situational awareness and cross-modal reasoning, pushing the boundaries of multimodal intelligence. This survey not only synthesizes existing research but also identifies critical challenges and future directions for building audio-native AGI systems capable of perceiving, understanding, and interacting through sound as naturally as humans do.

Keywords

Cite

@article{arxiv.2511.01299,
  title  = {Towards General Auditory Intelligence: Large Multimodal Models for Machine Listening and Speaking},
  author = {Siyin Wang and Zengrui Jin and Changli Tang and Qiujia Li and Bo Li and Chen Chen and Yuchen Hu and Wenyi Yu and Yixuan Li and Jimin Zhuang and Yudong Yang and Mingqiu Wang and Michael Han and Yifan Ding and Junwen Bai and Tom Ouyang and Shuo-yiin Chang and Xianzhao Chen and Xiaohai Tian and Jun Zhang and Lu Lu and Guangzhi Sun and Zhehuai Chen and Ji Wu and Bowen Zhou and Yuxuan Wang and Tara Sainath and Yonghui Wu and Chao Zhang},
  journal= {arXiv preprint arXiv:2511.01299},
  year   = {2025}
}

Comments

22 pages, 11 figures

R2 v1 2026-07-01T07:18:46.593Z