English
Related papers

Related papers: AudSemThinker: Enhancing Audio-Language Models thr…

200 papers

Recent advancements in large language models, multimodal large language models, and large audio language models (LALMs) have significantly improved their reasoning capabilities through reinforcement learning with rule-based rewards.…

Sound · Computer Science 2025-11-05 Shu Wu , Chenxing Li , Wenfu Wang , Hao Zhang , Hualei Wang , Meng Yu , Dong Yu

While large language models have demonstrated impressive reasoning abilities, their extension to the audio modality, particularly within large audio-language models (LALMs), remains underexplored. Addressing this gap requires a systematic…

Computation and Language · Computer Science 2025-09-23 Xingjian Diao , Chunhui Zhang , Keyi Kong , Weiyi Wu , Chiyu Ma , Zhongyu Ouyang , Peijun Qing , Soroush Vosoughi , Jiang Gui

Recent advancements in multimodal reasoning have largely overlooked the audio modality. We introduce Audio-Reasoner, a large-scale audio language model for deep reasoning in audio tasks. We meticulously curated a large-scale and diverse…

Sound · Computer Science 2025-09-23 Zhifei Xie , Mingbao Lin , Zihang Liu , Pengcheng Wu , Shuicheng Yan , Chunyan Miao

Audio-visual pre-trained models have gained substantial attention recently and demonstrated superior performance on various audio-visual tasks. This study investigates whether pre-trained audio-visual models demonstrate non-arbitrary…

Computation and Language · Computer Science 2024-11-13 Wei-Cheng Tseng , Yi-Jen Shih , David Harwath , Raymond Mooney

Recent Large Audio-Language Models (LALMs) have shown strong performance on various audio understanding tasks such as speech translation and Audio Q\&A. However, they exhibit significant limitations on challenging audio reasoning tasks in…

Computation and Language · Computer Science 2025-09-29 Zhen Xiong , Yujun Cai , Zhecheng Li , Junsong Yuan , Yiwei Wang

Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from those of text and vision. It is continuous, temporally dense,…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-21 Zhihan Guo , Wenqian Cui , Guan-Ting Lin , Daxin Tan , Jingyao Li , Qiyong Zheng , Dingdong Wang , Jing Xiong , Han Shi , Jiaya Jia , Irwin King

Even without directly hearing sounds, humans can effortlessly reason about auditory properties, such as pitch, loudness, or sound-source associations, drawing on auditory commonsense. In contrast, language models often lack this capability,…

Computation and Language · Computer Science 2026-01-29 Hyunjong Ok , Suho Yoo , Hyeonjun Kim , Jaeho Lee

Supervised learning methods can solve the given problem in the presence of a large set of labeled data. However, the acquisition of a dataset covering all the target classes typically requires manual labeling which is expensive and…

Sound · Computer Science 2022-06-13 Duygu Dogan , Huang Xie , Toni Heittola , Tuomas Virtanen

Large Audio-Language Models (LALMs) have made significant progress in audio understanding, yet they primarily operate as perception-and-answer systems without explicit reasoning processes. Existing methods for enhancing audio reasoning rely…

Sound · Computer Science 2026-04-21 Xiang He , Chenxing Li , Jinting Wang , Yan Rong , Tianxin Xie , Wenfu Wang , Li Liu , Dong Yu

(Part of the abstract) In this thesis, we investigate the use of unsupervised spoken term discovery in tackling this problem. Unsupervised spoken term discovery aims to discover topic-related terminologies in a speech without knowing the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-01 Man-Ling Sung

We study two foundational problems in audio language models: (1) how to design an audio tokenizer that can serve as an intermediate representation for both understanding and generation; and (2) how to build an audio foundation model that…

Sound · Computer Science 2026-02-12 Dongchao Yang , Yuanyuan Wang , Dading Chong , Songxiang Liu , Xixin Wu , Helen Meng

Multi-modal learning in the audio-language domain has seen significant advancements in recent years. However, audio-language learning faces challenges due to limited and lower-quality data compared to image-language tasks. Existing…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-10 David Xu

The ability of artificial intelligence (AI) systems to perceive and comprehend audio signals is crucial for many applications. Although significant progress has been made in this area since the development of AudioSet, most existing models…

Audio and Speech Processing · Electrical Eng. & Systems 2024-02-21 Yuan Gong , Hongyin Luo , Alexander H. Liu , Leonid Karlinsky , James Glass

Large Language Models (LLMs) have made significant progress in various downstream tasks, inspiring the development of Speech Understanding Language Models (SULMs) to enable comprehensive speech-based interactions. However, most advanced…

The text generation paradigm for audio tasks has opened new possibilities for unified audio understanding. However, existing models face significant challenges in achieving a comprehensive understanding across diverse audio types, such as…

Audio and Speech Processing · Electrical Eng. & Systems 2025-05-28 Ziqian Wang , Xianjun Xia , Xinfa Zhu , Lei Xie

In this paper, we present an audio analyzer assistant tool designed for a wide range of audio-based surveillance applications (This work is a part of our DEFAME FAKES and EUCINF projects). The proposed tool, refered to as Aud-Sur, comprises…

Audio and Speech Processing · Electrical Eng. & Systems 2025-04-01 Phat Lam , Lam Pham , Dat Tran , Alexander Schindler , Silvia Poletti , Marcel Hasenbalg , David Fischinger , Martin Boyer

We introduce AudioLM, a framework for high-quality audio generation with long-term consistency. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation…

Recent literature uses language to build foundation models for audio. These Audio-Language Models (ALMs) are trained on a vast number of audio-text pairs and show remarkable performance in tasks including Text-to-Audio Retrieval,…

Audio agents extend large audio-language models (LALMs) by decomposing audio questions into tool calls, intermediate evidence, and iterative reasoning steps. However, as LALMs become stronger, the key challenge shifts from enabling tool use…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-28 Yucheng Wang , Jing Peng , Hanqi Li , Chenghao Wang , Wenming Tu , Yu Xi , Zhaokai Sun , Kai Yu , Shuai Wang

Discrete audio representations are gaining traction in speech modeling due to their interpretability and compatibility with large language models, but are not always optimized for noisy or real-world environments. Building on existing works…

Computation and Language · Computer Science 2025-10-30 Shreyas Gopal , Ashutosh Anshul , Haoyang Li , Yue Heng Yeo , Hexin Liu , Eng Siong Chng
‹ Prev 1 2 3 10 Next ›