中文
相关论文

相关论文: MMAR: A Challenging Benchmark for Deep Reasoning i…

200 篇论文

Due to recent advancements in Large Audio-Language Models (LALMs) that demonstrate remarkable performance across a range of sound-, speech- and music-related tasks, there is a growing interest in proposing benchmarks to assess these models.…

音频与语音处理 · 电气工程与系统科学 2026-02-12 Jingru Lin , Chen Zhang , Tianrui Wang , Haizhou Li

Recent Large Audio Language Models (LALMs) excel in understanding but often lack transparent reasoning. To address this "black-box" limitation, we organized the Audio Reasoning Challenge at Interspeech 2026, the first shared task dedicated…

Large audio-language models (LALMs) have achieved near-human performance in sentence-level transcription and emotion recognition. However, existing evaluations focus mainly on surface-level perception, leaving the capacity of models for…

计算与语言 · 计算机科学 2025-08-05 Wanqi Yang , Yanda Li , Yunchao Wei , Meng Fang , Ling Chen

Large audio language models (ALMs) extend LLMs with auditory understanding. A common approach freezes the LLM and trains only an adapter on self-generated targets. However, this fails for reasoning LLMs (RLMs) whose built-in…

计算与语言 · 计算机科学 2026-03-11 Petr Grinberg , Hassan Shahmohammadi

With advancements in large audio-language models (LALMs), which enhance large language models (LLMs) with auditory capabilities, these models are expected to demonstrate universal proficiency across various auditory tasks. While numerous…

音频与语音处理 · 电气工程与系统科学 2026-04-28 Chih-Kai Yang , Neo S. Ho , Hung-yi Lee

Processing long-form audio is a major challenge for Large Audio Language models (LALMs). These models struggle with the quadratic cost of attention ($O(N^2)$) and with modeling long-range temporal dependencies. Existing audio benchmarks are…

Large Audio Language Models (LALMs) are increasingly capable of reasoning over audio. However, existing benchmarks provide limited coverage of reasoning in polyphonic audio, where multiple sound events co-occur and induce compositional…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Yuanjian Chen , Yang Xiao , Han Yin , Xubo Liu , Jinjie Huang , Ting Dang

The present benchmarks for testing the audio modality of multimodal large language models concentrate on testing various audio tasks such as speaker diarization or gender identification in isolation. Whether a multimodal model can answer…

Large Audio-Language Models (LALMs) have made significant progress in audio understanding, yet they primarily operate as perception-and-answer systems without explicit reasoning processes. Existing methods for enhancing audio reasoning rely…

声音 · 计算机科学 2026-04-21 Xiang He , Chenxing Li , Jinting Wang , Yan Rong , Tianxin Xie , Wenfu Wang , Li Liu , Dong Yu

The ability to reason from audio, including speech, environmental sounds, and music, is essential for AI agents to interact effectively in real-world scenarios. Existing benchmarks mainly focus on static or single-scene settings and English…

Perceiving and understanding non-speech sounds and non-verbal speech is essential to making decisions that help us interact with our surroundings. In this paper, we propose GAMA, a novel General-purpose Large Audio-Language Model (LALM)…

The ability to comprehend audio--which includes speech, non-speech sounds, and music--is crucial for AI agents to interact effectively with the world. We present MMAU, a novel benchmark designed to evaluate multimodal audio understanding…

音频与语音处理 · 电气工程与系统科学 2024-10-28 S Sakshi , Utkarsh Tyagi , Sonal Kumar , Ashish Seth , Ramaneswaran Selvakumar , Oriol Nieto , Ramani Duraiswami , Sreyan Ghosh , Dinesh Manocha

While large language models have demonstrated impressive reasoning abilities, their extension to the audio modality, particularly within large audio-language models (LALMs), remains underexplored. Addressing this gap requires a systematic…

计算与语言 · 计算机科学 2025-09-23 Xingjian Diao , Chunhui Zhang , Keyi Kong , Weiyi Wu , Chiyu Ma , Zhongyu Ouyang , Peijun Qing , Soroush Vosoughi , Jiang Gui

Large language models (LLMs) have advanced in text and vision, but their reasoning on audio remains limited. Most existing methods rely on dense audio embeddings, which are difficult to interpret and often fail on structured reasoning…

声音 · 计算机科学 2025-11-11 Termeh Taheri , Yinghao Ma , Emmanouil Benetos

Various audio-LLMs (ALLMs) have been explored recently for tackling different audio tasks simultaneously using a single, unified model. While existing evaluations of ALLMs primarily focus on single-audio tasks, real-world applications often…

声音 · 计算机科学 2024-11-07 Yiming Chen , Xianghu Yue , Xiaoxue Gao , Chen Zhang , Luis Fernando D'Haro , Robby T. Tan , Haizhou Li

Logical reasoning is a fundamental aspect of human intelligence and an essential capability for multimodal large language models (MLLMs). Despite the significant advancement in multimodal reasoning, existing benchmarks fail to…

人工智能 · 计算机科学 2025-05-28 Jiakang Yuan , Tianshuo Peng , Yilei Jiang , Yiting Lu , Renrui Zhang , Kaituo Feng , Chaoyou Fu , Tao Chen , Lei Bai , Bo Zhang , Xiangyu Yue

Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is largely generic (e.g., summarizing spoken content) and fails to…

计算与语言 · 计算机科学 2026-01-08 Yuwen Wang , Xinyuan Qian , Tian-Hao Zhang , Jiaran Gao , Yuchen Pan , Xin Wang , Zhou Pan , Chen Wei , Yiming Wang

Multimodal Large Language Models (MLLMs) have made substantial progress in recent years. However, their rigorous evaluation within specialized domains like finance is hindered by the absence of datasets characterized by professional-level…

人工智能 · 计算机科学 2025-11-25 Shuangyan Deng , Haizhou Peng , Jiachen Xu , Rui Mao , Ciprian Doru Giurcăneanu , Jiamou Liu

Multimodal Audio-Language Models (ALMs) can understand and reason over both audio and text. Typically, reasoning performance correlates with model size, with the best results achieved by models exceeding 8 billion parameters. However, no…

声音 · 计算机科学 2025-03-12 Soham Deshmukh , Satvik Dixit , Rita Singh , Bhiksha Raj

Reasoning has become a defining capability of modern foundation models, yet its development in the audio modality remains limited. Audio poses challenges that are distinct from those of text and vision. It is continuous, temporally dense,…

音频与语音处理 · 电气工程与系统科学 2026-05-21 Zhihan Guo , Wenqian Cui , Guan-Ting Lin , Daxin Tan , Jingyao Li , Qiyong Zheng , Dingdong Wang , Jing Xiong , Han Shi , Jiaya Jia , Irwin King
‹ 上一页 1 2 3 10 下一页 ›