中文
相关论文

相关论文: Seeing and Hearing Egocentric Actions: How Much Ca…

200 篇论文

Egocentric videos capture sequences of human activities from a first-person perspective and can provide rich multimodal signals. However, most current localization methods use third-person videos and only incorporate visual information. In…

计算机视觉与模式识别 · 计算机科学 2022-10-27 Merey Ramazanova , Victor Escorcia , Fabian Caba Heilbron , Chen Zhao , Bernard Ghanem

We propose a self-supervised method for learning representations based on spatial audio-visual correspondences in egocentric videos. Our method uses a masked auto-encoding framework to synthesize masked binaural (multi-channel) audio…

计算机视觉与模式识别 · 计算机科学 2024-05-07 Sagnik Majumder , Ziad Al-Halah , Kristen Grauman

Perceiving the world from both egocentric (first-person) and exocentric (third-person) perspectives is fundamental to human cognition, enabling rich and complementary understanding of dynamic environments. In recent years, allowing the…

计算机视觉与模式识别 · 计算机科学 2025-06-09 Yuping He , Yifei Huang , Guo Chen , Lidong Lu , Baoqi Pei , Jilan Xu , Tong Lu , Yoichi Sato

We study object interaction anticipation in egocentric videos. This task requires an understanding of the spatio-temporal context formed by past actions on objects, coined action context. We propose TransFusion, a multimodal…

计算机视觉与模式识别 · 计算机科学 2024-03-12 Razvan-George Pasca , Alexey Gavryushin , Muhammad Hamza , Yen-Ling Kuo , Kaichun Mo , Luc Van Gool , Otmar Hilliges , Xi Wang

Egocentric action anticipation consists in predicting a future action the camera wearer will perform from egocentric video. While the task has recently attracted the attention of the research community, current approaches assume that the…

计算机视觉与模式识别 · 计算机科学 2022-02-10 Ivan Rodin , Antonino Furnari , Dimitrios Mavroeidis , Giovanni Maria Farinella

Humans have the natural ability to recognize actions even if the objects involved in the action or the background are changed. Humans can abstract away the action from the appearance of the objects which is referred to as compositionality…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Ramanathan Rajendiran , Debaditya Roy , Basura Fernando

Wearable cameras allow to acquire images and videos from the user's perspective. These data can be processed to understand humans behavior. Despite human behavior analysis has been thoroughly investigated in third person vision, it is still…

计算机视觉与模式识别 · 计算机科学 2023-07-06 Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

The integration of information across multiple modalities and across time is a promising way to enhance the emotion recognition performance of affective systems. Much previous work has focused on instantaneous emotion recognition. The 2018…

图像与视频处理 · 电气工程与系统科学 2018-05-07 Didan Deng , Yuqian Zhou , Jimin Pi , Bertram E. Shi

As humans, we experience the world with all our senses or modalities (sound, sight, touch, smell, and taste). We use these modalities, particularly sight and touch, to convey and interpret specific meanings. Multimodal expressions are…

机器学习 · 计算机科学 2022-05-17 Anirudh Sundar , Larry Heck

First-person vision is gaining interest as it offers a unique viewpoint on people's interaction with objects, their attention, and even intention. However, progress in this challenging domain has been relatively slow due to the lack of…

Egocentric vision is an emerging field of computer vision that is characterized by the acquisition of images and video from the first person perspective. In this paper we address the challenge of egocentric human action recognition by…

计算机视觉与模式识别 · 计算机科学 2019-05-03 Georgios Kapidis , Ronald Poppe , Elsbeth van Dam , Lucas P. J. J. Noldus , Remco C. Veltkamp

This research aims to comprehensively explore building a multimodal foundation model for egocentric video understanding. To achieve this goal, we work on three fronts. First, as there is a lack of QA data for egocentric video understanding,…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Hanrong Ye , Haotian Zhang , Erik Daxberger , Lin Chen , Zongyu Lin , Yanghao Li , Bowen Zhang , Haoxuan You , Dan Xu , Zhe Gan , Jiasen Lu , Yinfei Yang

Multimodal video understanding is crucial for analyzing egocentric videos, where integrating multiple sensory signals significantly enhances action recognition and moment localization. However, practical applications often grapple with…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Merey Ramazanova , Alejandro Pardo , Humam Alwassel , Bernard Ghanem

Can conversational videos captured from multiple egocentric viewpoints reveal the map of a scene in a cost-efficient way? We seek to answer this question by proposing a new problem: efficiently building the map of a previously unseen 3D…

计算机视觉与模式识别 · 计算机科学 2023-04-24 Sagnik Majumder , Hao Jiang , Pierre Moulon , Ethan Henderson , Paul Calamia , Kristen Grauman , Vamsi Krishna Ithapu

With the availability of egocentric 3D hand-object interaction datasets, there is increasing interest in developing unified models for hand-object pose estimation and action recognition. However, existing methods still struggle to recognise…

计算机视觉与模式识别 · 计算机科学 2025-01-14 Tze Ho Elden Tse , Runyang Feng , Linfang Zheng , Jiho Park , Yixing Gao , Jihie Kim , Ales Leonardis , Hyung Jin Chang

Anticipating future actions is a highly challenging task due to the diversity and scale of potential future actions; yet, information from different modalities help narrow down plausible action choices. Each modality can provide diverse and…

计算机视觉与模式识别 · 计算机科学 2024-08-30 Apoorva Beedu , Harish Haresamudram , Karan Samel , Irfan Essa

As the prevalence of wearable devices, learning egocentric motions becomes essential to develop contextual AI. In this work, we present EgoLM, a versatile framework that tracks and understands egocentric motions from multi-modal inputs,…

计算机视觉与模式识别 · 计算机科学 2024-09-27 Fangzhou Hong , Vladimir Guzov , Hyo Jin Kim , Yuting Ye , Richard Newcombe , Ziwei Liu , Lingni Ma

In this paper, we present our solutions for emotion recognition in the sub-challenges of Multimodal Emotion Recognition Challenge (MER2024). To mitigate the modal competition issue between audio and text, we adopt an early fusion strategy…

多媒体 · 计算机科学 2024-10-01 Mengying Ge , Mingyang Li , Dongkai Tang , Pengbo Li , Kuo Liu , Shuhao Deng , Songbai Pu , Long Liu , Yang Song , Tao Zhang

We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and stationary cameras, we recorded synchronized gaze, speech, and…

计算机视觉与模式识别 · 计算机科学 2025-10-29 Anna Deichler , Jonas Beskow

As AI agents increasingly operate in open, real-world environments, they require a deep synergy of multimodal perception, tool invocation with multi-hop reasoning, and dynamic interaction with users. However, existing benchmarks fail to…

人工智能 · 计算机科学 2026-05-28 Yunqi Liu , Tong Niu , Zitong Wang , Zhenlong Dai , Yuqi Qing , Weiqiang Wang , Jian Liu