English
Related papers

Related papers: Seeing and Hearing Egocentric Actions: How Much Ca…

200 papers

Humans excel at multisensory perception and can often recognise object properties from the sound of their interactions. Inspired by this, we propose the novel task of Collision Sound Source Segmentation (CS3), where we aim to segment the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-21 Kranti Kumar Parida , Omar Emara , Hazel Doughty , Dima Damen

Contact-rich manipulation requires reliable estimation of extrinsic contacts-the interactions between a grasped object and its environment which provide essential contextual information for planning, control, and policy learning. However,…

Robotics · Computer Science 2026-02-03 Zhengtong Xu , Yuki Shirai

Multimodal Large Language Models (MLLMs) have recently made rapid progress toward unified Omni models that integrate vision, language, and audio. However, existing environments largely focus on 2D or 3D visual context and vision-language…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yurui Dong , Ziyue Wang , Shuyun Lu , Dairu Liu , Xuechen Liu , Fuwen Luo , Peng Li , Yang Liu

Egocentric video understanding is inherently complex due to the dynamic 4D nature of the environment, where camera motion and object displacements necessitate a continuous re-evaluation of spatial relations. In this work, we target a suite…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Fangrui Zhu , Yunfeng Xi , Jianmo Ni , Mu Cai , Boqing Gong , Long Zhao , Chen Qu , Ian Miao , Yi Li , Cheng Zhong , Huaizu Jiang , Shwetak Patel

Self-supervised learning has attracted plenty of recent research interest. However, most works for self-supervision in speech are typically unimodal and there has been limited work that studies the interaction between audio and visual…

Audio and Speech Processing · Electrical Eng. & Systems 2021-03-19 Abhinav Shukla , Stavros Petridis , Maja Pantic

The recent explosion of interest in multimodal applications has resulted in a wide selection of datasets and methods for representing and integrating information from different modalities. Despite these empirical advances, there remain…

Human perception of the empirical world involves recognizing the diverse appearances, or 'modalities', of underlying objects. Despite the longstanding consideration of this perspective in philosophy and cognitive science, the study of…

Machine Learning · Computer Science 2023-12-19 Zhou Lu

Existing top-performance autonomous driving systems typically rely on the multi-modal fusion strategy for reliable scene understanding. This design is however fundamentally restricted due to overlooking the modality-specific strengths and…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Zeyu Yang , Nan Song , Wei Li , Xiatian Zhu , Li Zhang , Philip H. S. Torr

Videos are more well-organized curated data sources for visual concept learning than images. Unlike the 2-dimensional images which only involve the spatial information, the additional temporal dimension bridges and synchronizes multiple…

Computer Vision and Pattern Recognition · Computer Science 2022-05-13 Keren Ye , Adriana Kovashka

This paper presents a framework for recognition of human activity from egocentric video and eye tracking data obtained from a head-mounted eye tracker. Three channels of information such as eye movement, ego-motion, and visual features are…

Computer Vision and Pattern Recognition · Computer Science 2018-05-21 Anjith George , Aurobinda Routray

In this report, we describe our approach to egocentric video object segmentation. Our method combines large-scale visual pretraining from SAM2 with depth-based geometric cues to handle complex scenes and long-term tracking. By integrating…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Mingqi Gao , Haoran Duan , Tianlu Zhang , Jungong Han

In this paper, we tackle the problem of Egocentric Human-Object Interaction (EHOI) detection in an industrial setting. To overcome the lack of public datasets in this context, we propose a pipeline and a tool for generating synthetic images…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Rosario Leonardi , Francesco Ragusa , Antonino Furnari , Giovanni Maria Farinella

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the…

Computer Vision and Pattern Recognition · Computer Science 2023-11-13 Tushar Nagarajan , Santhosh Kumar Ramakrishnan , Ruta Desai , James Hillis , Kristen Grauman

In this paper, we propose a method to jointly determine the status of hand-object interaction. This is crucial for egocentric human activity understanding and interaction. From a computer vision perspective, we believe that determining…

Computer Vision and Pattern Recognition · Computer Science 2022-11-17 Yao Lu , Yanan Liu

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Ashish Seth , Xinhao Mei , Changsheng Zhao , Varun Nagaraja , Ernie Chang , Gregory P. Meyer , Gael Le Lan , Yunyang Xiong , Vikas Chandra , Yangyang Shi , Dinesh Manocha , Zhipeng Cai

Egocentric videos provide a unique perspective into individuals' daily experiences, yet their unstructured nature presents challenges for perception. In this paper, we introduce AMEGO, a novel approach aimed at enhancing the comprehension…

Computer Vision and Pattern Recognition · Computer Science 2024-09-18 Gabriele Goletto , Tushar Nagarajan , Giuseppe Averta , Dima Damen

Although there has been significant research in egocentric action recognition, most methods and tasks, including EPIC-KITCHENS, suppose a fixed set of action classes. Fixed-set classification is useful for benchmarking methods, but is often…

Computer Vision and Pattern Recognition · Computer Science 2020-06-23 Tyler R. Scott , Michael Shvartsman , Karl Ridgeway

In the human activity recognition research area, prior studies predominantly concentrate on leveraging advanced algorithms on public datasets to enhance recognition performance, little attention has been paid to executing real-time kitchen…

Signal Processing · Electrical Eng. & Systems 2024-09-11 Mengxi Liu , Sungho Suh , Juan Felipe Vargas , Bo Zhou , Agnes Grünerbl , Paul Lukowicz

Multimodal Language Analysis is a demanding area of research, since it is associated with two requirements: combining different modalities and capturing temporal information. During the last years, several works have been proposed in the…

Computation and Language · Computer Science 2022-01-10 Panagiotis Koromilas , Theodoros Giannakopoulos

Vision and voice are two vital keys for agents' interaction and learning. In this paper, we present a novel indoor navigation model called Memory Vision-Voice Indoor Navigation (MVV-IN), which receives voice commands and analyzes multimodal…

Computer Vision and Pattern Recognition · Computer Science 2020-09-02 Liqi Yan , Dongfang Liu , Yaoxian Song , Changbin Yu