中文
相关论文

相关论文: When Audio and Text Disagree: Revealing Text Bias …

200 篇论文

Recently, Large Audio Language Models (LALMs) have progressed rapidly, demonstrating their strong efficacy in universal audio understanding through cross-modal integration. To evaluate LALMs' audio understanding performance, researchers…

声音 · 计算机科学 2026-02-16 Han Yin , Jung-Woo Choi

Large Audio-Language Models (LALMs) as judges have emerged as a prominent approach for evaluating speech generation quality, yet their ability to assess speaker consistency across multi-turn dialogues remains unexplored. We present…

计算与语言 · 计算机科学 2026-04-21 Jonggeun Lee , Junseong Pyo , Gyuhyeon Seo , Yohan Jo

Large Audio-Language Models (LALMs) have demonstrated strong performance in audio understanding and generation. Yet, our extensive benchmarking reveals that their behavior is largely generic (e.g., summarizing spoken content) and fails to…

计算与语言 · 计算机科学 2026-01-08 Yuwen Wang , Xinyuan Qian , Tian-Hao Zhang , Jiaran Gao , Yuchen Pan , Xin Wang , Zhou Pan , Chen Wei , Yiming Wang

Large audio-language models (LALMs) unify speech and text processing, but their robustness in noisy real-world settings remains underexplored. We investigate how irrelevant audio, such as silence, synthetic noise, and environmental sounds,…

声音 · 计算机科学 2026-04-28 Chen-An Li , Tzu-Han Lin , Hung-yi Lee

Recently, instruction-following audio-language models have received broad attention for human-audio interaction. However, the absence of benchmarks capable of evaluating audio-centric interaction capabilities has impeded advancements in…

音频与语音处理 · 电气工程与系统科学 2024-07-29 Qian Yang , Jin Xu , Wenrui Liu , Yunfei Chu , Ziyue Jiang , Xiaohuan Zhou , Yichong Leng , Yuanjun Lv , Zhou Zhao , Chang Zhou , Jingren Zhou

The popular success of text-based large language models (LLM) has streamlined the attention of the multimodal community to combine other modalities like vision and audio along with text to achieve similar multimodal capabilities. In this…

计算与语言 · 计算机科学 2025-05-20 Debarpan Bhattacharya , Apoorva Kulkarni , Sriram Ganapathy

How does textual representation of audio relate to the Large Language Model's (LLMs) learning about the audio world? This research investigates the extent to which LLMs can be prompted to generate audio, despite their primary training in…

Audio Large Language Models (Audio LLMs) enable human-like conversation about music, yet it is unclear if they are truly listening to the audio or just using textual reasoning, as recent benchmarks suggest. This paper investigates this…

机器学习 · 计算机科学 2026-05-18 Giovana Morais , Magdalena Fuentes

Large Audio Language Models (LALMs) have emerged as powerful tools for speech-related tasks but remain underexplored for fine-tuning, especially with limited speech data. To bridge this gap, we systematically examine how different…

声音 · 计算机科学 2026-01-22 Youngwon Choi , Jaeyoon Jung , Hyeonyu Kim , Huu-Kim Nguyen , Hwayeon Kim

Can Multimodal Large Language Models (MLLMs) discern confused objects that are visually present but audio-absent? To study this, we introduce a new benchmark, AV-ConfuseBench, which simulates an ``Audio-Visual Confusion'' scene by modifying…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Qilang Ye , Wei Zeng , Meng Liu , Jie Zhang , Yupeng Hu , Zitong Yu , Yu Zhou

Automated Audio Captioning aims to describe the semantic content of input audio. Recent works have employed large language models (LLMs) as a text decoder to leverage their reasoning capabilities. However, prior approaches that project…

声音 · 计算机科学 2026-03-17 Hyeongkeun Lee , Jongmin Choi , KiHyun Nam , Joon Son Chung

Large language models (LLMs) have been widely used as knowledge backbones of Large Audio Language Models (LALMs), yet how much auditory knowledge they encode through text-only pre-training and how this affects downstream performance remains…

When audio and text conflict, speech-enabled language models follow text far more often than they do when arbitrating between two conflicting text sources, even under explicit instructions to trust the audio. We introduce ALME (Audio-LLM…

计算与语言 · 计算机科学 2026-03-25 Jayadev Billa

While Large Audio-Language Models (LALMs) have been shown to exhibit degraded instruction-following capabilities, their ability to infer task patterns from in-context examples under audio conditioning remains unstudied. To address this gap,…

Large Audio-Language Models (LALMs) often suffer from audio-textual attention imbalance, prioritizing text over acoustic information, particularly in the multi-modal fusion layers of the Transformer architecture. This bias hinders their…

声音 · 计算机科学 2025-09-24 Junyu Wang , Ziyang Ma , Zhengding Luo , Tianrui Wang , Meng Ge , Xiaobao Wang , Longbiao Wang

Large Audio Language Models (LALMs), which couple acoustic perception with large language models (LLMs) to extract and understand diverse information from audio, have attracted intense interest from both academic and industrial communities.…

声音 · 计算机科学 2025-10-28 Bohan Li , Wenbin Huang , Yuhang Qiu , Yiwei Guo , Hankun Wang , Zhihan Li , Jing Peng , Ziyang Ma , Xie Chen , Kai Yu

Understanding the internal mechanisms of large audio-language models (LALMs) is crucial for interpreting their behavior and improving performance. This work presents the first in-depth analysis of how LALMs internally perceive and recognize…

计算与语言 · 计算机科学 2025-08-26 Chih-Kai Yang , Neo Ho , Yi-Jyun Lee , Hung-yi Lee

Large Audio Language Models (LALMs) are increasingly applied to audio understanding and multimodal reasoning, yet their ability to locate when events occur remains underexplored. We present the first systematic study of temporal bias in…

计算与语言 · 计算机科学 2025-10-15 Jiayu Yao , Shenghua Liu , Yiwei Wang , Rundong Cheng , Lingrui Mei , Baolong Bi , Zhen Xiong , Xueqi Cheng

As large language models transition from text-based interfaces to audio interactions in clinical settings, they might introduce new vulnerabilities through paralinguistic cues in audio. We evaluated these models on 170 clinical cases, each…

计算与语言 · 计算机科学 2025-11-11 Zhi Rui Tam , Yun-Nung Chen

Despite remarkable advancements in Multimodal Large Language Models (MLLMs), a fundamental question remains: are MLLMs robust to contradicting modalities? To rigorously study this, we introduce MMA-Bench comprising videos and tasks that…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Tianle Chen , Chaitanya Chakka , Arjun Reddy Akula , Xavier Thomas , Deepti Ghadiyaram
‹ 上一页 1 2 3 10 下一页 ›