English
Related papers

Related papers: Exploring Audio Hallucination in Egocentric Video …

200 papers

Multimodal Large Language Models (MLLMs) have demonstrated remarkable video reasoning capabilities across diverse tasks. However, their ability to understand human intent at a fine-grained level in egocentric videos remains largely…

Computer Vision and Pattern Recognition · Computer Science 2026-03-13 Ye Pan , Chi Kit Wong , Yuanhuiyi Lyu , Hanqian Li , Jiahao Huo , Jiacheng Chen , Lutao Jiang , Xu Zheng , Xuming Hu

Despite strong performance in audio perception tasks, large audio-language models (AudioLLMs) remain opaque to interpretation. A major factor behind this lack of interpretability is that individual neurons in these models frequently…

Sound · Computer Science 2026-02-27 Townim Faisal Chowdhury , Ta Duc Huy , Siqi Pan , Jeremy Stoddard , Zhibin Liao

Hallucination has been a major problem for large language models and remains a critical challenge when it comes to multimodality in which vision-language models (VLMs) have to deal with not just textual but also visual inputs. Despite rapid…

Computer Vision and Pattern Recognition · Computer Science 2024-07-23 Zhecan Wang , Garrett Bingham , Adams Yu , Quoc Le , Thang Luong , Golnaz Ghiasi

Today's advanced driver assistance systems (ADAS), like adaptive cruise control or rear collision warning, are finding broader adoption across vehicle classes. Integrating such advanced, multimodal Large Language Models (LLMs) on board a…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Malsha Ashani Mahawatta Dona , Beatriz Cabrero-Daniel , Yinan Yu , Christian Berger

Large vision-language models (LVLMs) are prone to hallucinations, where certain contextual cues in an image can trigger the language module to produce overconfident and incorrect reasoning about abnormal or hypothetical objects. While some…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Xiyang Wu , Tianrui Guan , Dianqi Li , Shuaiyi Huang , Xiaoyu Liu , Xijun Wang , Ruiqi Xian , Abhinav Shrivastava , Furong Huang , Jordan Lee Boyd-Graber , Tianyi Zhou , Dinesh Manocha

Humans excel at multisensory perception and can often recognise object properties from the sound of their interactions. Inspired by this, we propose the novel task of Collision Sound Source Segmentation (CS3), where we aim to segment the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Kranti Kumar Parida , Omar Emara , Hazel Doughty , Dima Damen

Multimodal large language models (MLLMs) are expected to jointly interpret vision, audio, and language, yet existing video benchmarks rarely assess fine-grained reasoning about human speech. Many tasks remain visually solvable or only…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Le Thien Phuc Nguyen , Zhuoran Yu , Samuel Low Yu Hang , Subin An , Jeongik Lee , Yohan Ban , SeungEun Chung , Thanh-Huy Nguyen , JuWan Maeng , Soochahn Lee , Yong Jae Lee

We present the first systematic analysis of multimodal large language models (MLLMs) in personalized question-answering requiring ego-grounding - the ability to understand the camera-wearer in egocentric videos. To this end, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Junbin Xiao , Shenglang Zhang , Pengxiang Zhu , Angela Yao

Large Audio Language Models (LALMs) achieve strong performance on audio-language tasks; however, their reliability in real-world settings remains underexplored. We introduce Audio Hallucination Attacks (AHA), an attack suite called…

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sudha Krishnamurthy

We introduce a state-of-the-art audio-visual on-screen sound separation system which is capable of learning to separate sounds and associate them with on-screen objects by looking at in-the-wild videos. We identify limitations of previous…

Sound · Computer Science 2021-10-15 Efthymios Tzinis , Scott Wisdom , Tal Remez , John R. Hershey

Egocentric activity recognition in first-person videos has an increasing importance with a variety of applications such as lifelogging, summarization, assisted-living and activity tracking. Existing methods for this task are based on…

Computer Vision and Pattern Recognition · Computer Science 2020-05-01 Mehmet Ali Arabacı , Fatih Özkan , Elif Surer , Peter Jančovič , Alptekin Temizel

Recent audio-aware large language models (ALLMs) have demonstrated strong capabilities across diverse audio understanding and reasoning tasks, but they still frequently produce hallucinated or overly confident outputs. While uncertainty…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-29 Chun-Yi Kuan , Wei-Ping Huang , Hung-yi Lee

Wearable devices such as AI glasses are transforming voice assistants into always-available, hands-free collaborators that integrate seamlessly with daily life, but they also introduce challenges like egocentric audio affected by motion and…

The hallucination problem of Large Language Models (LLMs) significantly limits their reliability and trustworthiness. Humans have a self-awareness process that allows us to recognize what we don't know when faced with queries. Inspired by…

Computation and Language · Computer Science 2024-10-01 Ziwei Ji , Delong Chen , Etsuko Ishii , Samuel Cahyawijaya , Yejin Bang , Bryan Wilie , Pascale Fung

Humans can intuitively infer sounds from silent videos, but whether multimodal large language models can perform modal-mismatch reasoning without accessing target modalities remains relatively unexplored. Current…

Multimedia · Computer Science 2025-05-29 Yong Ren , Chenxing Li , Le Xu , Hao Gu , Duzhen Zhang , Yujie Chen , Manjie Xu , Ruibo Fu , Shan Yang , Dong Yu

Although Large Audio-Language Models (LALMs) deliver state-of-the-art (SOTA) performance, they frequently suffer from hallucinations, e.g. generating text not grounded in the audio input. We analyze these grounding failures and identify a…

Egocentric video understanding requires procedural reasoning under partial observability and continuously shifting viewpoints. Current multimodal large language models (MLLMs) struggle with this setting, often generating plausible but…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yogesh Kulkarni , Pooyan Fazli

Sounds are an important source of information on our daily interactions with objects. For instance, a significant amount of people can discern the temperature of water that it is being poured just by using the sense of hearing. However,…

Computer Vision and Pattern Recognition · Computer Science 2019-06-04 Alejandro Cartas , Jordi Luque , Petia Radeva , Carlos Segura , Mariella Dimiccoli

Recent advancements in large multimodal models (LMMs) have significantly enhanced performance across diverse tasks, with ongoing efforts to further integrate additional modalities such as video and audio. However, most existing LMMs remain…

Computer Vision and Pattern Recognition · Computer Science 2024-10-17 Sicong Leng , Yun Xing , Zesen Cheng , Yang Zhou , Hang Zhang , Xin Li , Deli Zhao , Shijian Lu , Chunyan Miao , Lidong Bing