中文
相关论文

相关论文: EgoAVU: Egocentric Audio-Visual Understanding

200 篇论文

Large Audio Language Models (LALMs) still struggle in complex acoustic scenes because they often fail to preserve task-relevant acoustic evidence before reasoning begins. We identify this error pattern as the evidence bottleneck:…

声音 · 计算机科学 2026-05-29 Xinyuan Xie , Shunian Chen , Zhiheng Liu , Yuhao Zhang , Zhiqiang Lv , Liyin Liang , Benyou Wang

Understanding human behavior requires measuring behavioral actions. Due to its complexity, behavior is best mapped onto a rich, semantic structure such as language. Emerging multimodal large language models (MLLMs) are promising candidates,…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Haozhe Qi , Shaokai Ye , Alexander Mathis , Mackenzie W. Mathis

Recent Video Large Language Models (Video-LLMs) have demonstrated strong capabilities in video reasoning through reinforcement learning (RL). However, existing RL pipelines rely heavily on human-annotated tasks and solutions, making them…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Shiqi Huang , Ziyue Wang , Zhongrong Zuo , Han Qiu , Qi She , Bihan Wen

TTM (Talking to Me) task is a pivotal component in understanding human social interactions, aiming to determine who is engaged in conversation with the camera-wearer. Traditional models often face challenges in real-world scenarios due to…

多媒体 · 计算机科学 2026-03-20 Xinyuan Qian , Xinjia Zhu , Alessio Brutti , Dong Liang

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for the EPIC-KITCHENS-100 Multi-Instance Retrieval (MIR) challenge. Especially, we exploit the recently released Ego4D dataset…

Recent advances in video-large language models (Video-LLMs) have led to significant progress in video understanding. Current preference optimization methods often rely on proprietary APIs or human-annotated captions to generate preference…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Yogesh Kulkarni , Pooyan Fazli

Electroencephalography (EEG) interpretation using multimodal large language models (MLLMs) offers a novel approach for analyzing brain signals. However, the complex nature of brain activity introduces critical challenges: EEG signals…

信号处理 · 电气工程与系统科学 2025-10-02 Ziyi Zeng , Zhenyang Cai , Yixi Cai , Xidong Wang , Junying Chen , Rongsheng Wang , Yipeng Liu , Siqi Cai , Benyou Wang , Zhiguo Zhang , Haizhou Li

Large multimodal models (LMMs) have recently emerged as a powerful tool for long video understanding (LVU), prompting the development of standardized LVU benchmarks to evaluate their performance. However, our investigation reveals a rather…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Wentao Ma , Weiming Ren , Yiming Jia , Zhuofeng Li , Ping Nie , Ge Zhang , Wenhu Chen

Recent advances in multimodal large language models (MLLMs) have demonstrated substantial potential in video understanding. However, existing benchmarks fail to comprehensively evaluate synergistic reasoning capabilities across audio and…

As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. However, it remains challenging due to the inherent complexity…

计算机视觉与模式识别 · 计算机科学 2026-01-06 Sungjune Park , Hongda Mao , Qingshuang Chen , Yong Man Ro , Yelin Kim

Ultra-long egocentric videos spanning multiple days present significant challenges for video understanding. Existing approaches still rely on fragmented local processing and limited temporal modeling, restricting their ability to reason…

计算机视觉与模式识别 · 计算机科学 2026-03-02 Shitong Sun , Ke Han , Yukai Huang , Weitong Cai , Jifei Song

In egocentric videos, actions occur in quick succession. We capitalise on the action's temporal context and propose a method that learns to attend to surrounding actions in order to improve recognition performance. To incorporate the…

计算机视觉与模式识别 · 计算机科学 2021-11-02 Evangelos Kazakos , Jaesung Huh , Arsha Nagrani , Andrew Zisserman , Dima Damen

Visual Emotion Analysis (VEA) aims to bridge the affective gap between visual content and human emotional responses. Despite its promise, progress in this field remains limited by the lack of open-source and interpretable datasets. Most…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Yijie Guo , Dexiang Hong , Weidong Chen , Zihan She , Cheng Ye , Xiaojun Chang , Zhendong Mao

Vision Language Models (VLMs) have achieved strong performance across diverse video understanding tasks. However, their viewpoint invariant training limits their ability to understand egocentric properties (e.g., human object interactions)…

计算机视觉与模式识别 · 计算机科学 2025-12-17 Dominick Reilly , Manish Kumar Govind , Le Xue , Srijan Das

Recently, multimodal large language models (MLLMs), such as GPT-4o, Gemini 1.5 Pro, and Reka Core, have expanded their capabilities to include vision and audio modalities. While these models demonstrate impressive performance across a wide…

计算机视觉与模式识别 · 计算机科学 2024-12-04 Kaixiong Gong , Kaituo Feng , Bohao Li , Yibing Wang , Mofan Cheng , Shijia Yang , Jiaming Han , Benyou Wang , Yutong Bai , Zhuoran Yang , Xiangyu Yue

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

Multimodal large language models (MLLMs) have shown promising advancements in general visual and language understanding. However, the representation of multimodal information using MLLMs remains largely unexplored. In this work, we…

计算与语言 · 计算机科学 2024-07-18 Ting Jiang , Minghui Song , Zihan Zhang , Haizhen Huang , Weiwei Deng , Feng Sun , Qi Zhang , Deqing Wang , Fuzhen Zhuang

Audio Large Language Models (AudioLLMs) have achieved strong results in semantic tasks like speech recognition and translation, but remain limited in modeling paralinguistic cues such as emotion. Existing approaches often treat emotion…

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging…

机器学习 · 计算机科学 2026-04-10 Qiance Tang , Ziqi Wang , Jieyu Lin , Ziyun Li , Barbara De Salvo , Sai Qian Zhang

We present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive,…

计算机视觉与模式识别 · 计算机科学 2024-11-08 Keshigeyan Chandrasegaran , Agrim Gupta , Lea M. Hadzic , Taran Kota , Jimming He , Cristóbal Eyzaguirre , Zane Durante , Manling Li , Jiajun Wu , Li Fei-Fei