SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations
Abstract
Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder models are known to learn general-purpose and transferable representations, we investigate whether general-purpose SSL audio representations can serve as an effective foundation for ALLMs. We present SALMONN-2, an ALLM built upon a unified SSL encoder. To better exploit the hierarchical representations learned by SSL encoders, we propose a multi-layer feature fusion (MLF) adapter that aggregates information from all encoder layers before projecting them into the language model. Beyond conventional audio understanding tasks, we further explore multimodal in-context learning (MICL) in ALLMs and study how this capability can be acquired through contextual biasing training. Experimental results show that a general-purpose SSL encoder achieves performance comparable to, or better than, specialised supervised audio encoders while providing a more balanced capability across speech, audio, music and paralinguistic tasks. SALMONN-2 further achieves state-of-the-art performance among comparable-scale open-weight models on ALLM understanding benchmarks, obtaining the best results on MMAU-Pro, MMAR and MMSU. We also show that MICL does not emerge naturally in ALLMs, but can be effectively acquired through targeted contextual biasing training.
Cite
@article{arxiv.2607.17079,
title = {SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations},
author = {Xiaoyu Yang and Xuenan Xu and Wenyi Yu and Siyin Wang and Changli Tang and Terumi Chiba and Siyuan Hou and Ziyang Zhang and Wen Wu and Baoxiang Li and Guangzhi Sun and Chao Zhang and Philip Woodland},
journal= {arXiv preprint arXiv:2607.17079},
year = {2026}
}
Comments
Disclaimer: This work has been submitted to the IEEE for possible publication