English
Related papers

Related papers: The Audio-Visual Conversational Graph: From an Ego…

200 papers

Understanding egocentric videos plays a vital role for embodied intelligence. Recent multi-modal large language models (MLLMs) can accept both visual and audio inputs. However, due to the challenge of obtaining text labels with coherent…

Computer Vision and Pattern Recognition · Computer Science 2026-02-09 Ashish Seth , Xinhao Mei , Changsheng Zhao , Varun Nagaraja , Ernie Chang , Gregory P. Meyer , Gael Le Lan , Yunyang Xiong , Vikas Chandra , Yangyang Shi , Dinesh Manocha , Zhipeng Cai

Thanks to the availability and increasing popularity of Egocentric cameras such as GoPro cameras, glasses, and etc. we have been provided with a plethora of videos captured from the first person perspective. Surveillance cameras and…

Computer Vision and Pattern Recognition · Computer Science 2016-09-15 Shervin Ardeshir , Ali Borji

Egocentric videos offer fine-grained information for high-fidelity modeling of human behaviors. Hands and interacting objects are one crucial aspect of understanding a viewer's behaviors and intentions. We provide a labeled dataset…

Computer Vision and Pattern Recognition · Computer Science 2022-08-09 Lingzhi Zhang , Shenghao Zhou , Simon Stent , Jianbo Shi

Predicting when to initiate speech in real-world environments remains a fundamental challenge for conversational agents. We introduce EgoSpeak, a novel framework for real-time speech initiation prediction in egocentric streaming video. By…

Computer Vision and Pattern Recognition · Computer Science 2025-02-24 Junhyeok Kim , Min Soo Kim , Jiwan Chung , Jungbin Cho , Jisoo Kim , Sungwoong Kim , Gyeongbo Sim , Youngjae Yu

Automatically describing video, or video captioning, has been widely studied in the multimedia field. This paper proposes a new task of sensor-augmented egocentric-video captioning, a newly constructed dataset for it called MMAC Captions,…

Computer Vision and Pattern Recognition · Computer Science 2021-09-08 Katsuyuki Nakamura , Hiroki Ohashi , Mitsuhiro Okada

In this paper we propose a new framework to categorize social interactions in egocentric videos, we named InteractionGCN. Our method extracts patterns of relational and non-relational cues at the frame level and uses them to build a…

Computer Vision and Pattern Recognition · Computer Science 2021-06-09 Simone Felicioni , Mariella Dimiccoli

Anticipating human actions is an important task that needs to be addressed for the development of reliable intelligent agents, such as self-driving cars or robot assistants. While the ability to make future predictions with high accuracy is…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Olga Zatsarynna , Yazan Abu Farha , Juergen Gall

Recent advances in conversational AI have been substantial, but developing real-time systems for perceptual task guidance remains challenging. These systems must provide interactive, proactive assistance based on streaming visual inputs,…

Artificial Intelligence · Computer Science 2025-06-09 Yichi Zhang , Xin Luna Dong , Zhaojiang Lin , Andrea Madotto , Anuj Kumar , Babak Damavandi , Joyce Chai , Seungwhan Moon

We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and the future contact points on the next active object (i.e.,…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Shaowei Liu , Subarna Tripathi , Somdeb Majumdar , Xiaolong Wang

We present Egocentric Action Scene Graphs (EASGs), a new representation for long-form understanding of egocentric videos. EASGs extend standard manually-annotated representations of egocentric videos, such as verb-noun action labels, by…

Computer Vision and Pattern Recognition · Computer Science 2023-12-07 Ivan Rodin , Antonino Furnari , Kyle Min , Subarna Tripathi , Giovanni Maria Farinella

Predicting future human behavior from egocentric videos is a challenging but critical task for human intention understanding. Existing methods for forecasting 2D hand positions rely on visual representations and mainly focus on hand-object…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Masashi Hatano , Ryo Hachiuma , Hideo Saito

AI personal assistants, deployed through robots or wearables, require embodied understanding to collaborate effectively with humans. However, current Multimodal Large Language Models (MLLMs) primarily focus on third-person (exocentric)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Haoyu Zhang , Qiaohui Chu , Meng Liu , Haoxiang Shi , Yaowei Wang , Liqiang Nie

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Daniele Materia , Francesco Ragusa , Giovanni Maria Farinella

Joint visual attention (JVA) provides informative cues on human behavior during social interactions. The ubiquity of egocentric eye-trackers and large-scale datasets on everyday interactions offer research opportunities in identifying JVA…

Human-Computer Interaction · Computer Science 2025-09-17 Kumushini Thennakoon , Yasasi Abeysinghe , Bhanuka Mahanama , Vikas Ashok , Sampath Jayarathna

Wearable collaborative robots stand to assist human wearers who need fall prevention assistance or wear exoskeletons. Such a robot needs to be able to constantly adapt to the surrounding scene based on egocentric vision, and predict the ego…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Weizhuo Wang , C. Karen Liu , Monroe Kennedy

We present EMBED (Egocentric Models Built with Exocentric Data), a method designed to transform exocentric video-language data for egocentric video representation learning. Large-scale exocentric data covers diverse activities with…

Computer Vision and Pattern Recognition · Computer Science 2024-08-08 Zi-Yi Dou , Xitong Yang , Tushar Nagarajan , Huiyu Wang , Jing Huang , Nanyun Peng , Kris Kitani , Fu-Jen Chu

We introduce EASG-Bench, a question-answering benchmark for egocentric videos where the question-answering pairs are created from spatio-temporally grounded dynamic scene graphs capturing intricate relationships among actors, actions, and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Ivan Rodin , Tz-Ying Wu , Kyle Min , Sharath Nittur Sridhar , Antonino Furnari , Subarna Tripathi , Giovanni Maria Farinella

Egocentric gaze anticipation serves as a key building block for the emerging capability of Augmented Reality. Notably, gaze behavior is driven by both visual cues and audio signals during daily activities. Motivated by this observation, we…

Computer Vision and Pattern Recognition · Computer Science 2024-03-25 Bolin Lai , Fiona Ryan , Wenqi Jia , Miao Liu , James M. Rehg

Human-object interaction is one of the most important visual cues and we propose a novel way to represent human-object interactions for egocentric action anticipation. We propose a novel transformer variant to model interactions by…

Computer Vision and Pattern Recognition · Computer Science 2024-01-12 Debaditya Roy , Ramanathan Rajendiran , Basura Fernando

Our interaction with the world is an inherently multimodal experience. However, the understanding of human-to-object interactions has historically been addressed focusing on a single modality. In particular, a limited number of works have…

Computer Vision and Pattern Recognition · Computer Science 2019-10-16 Alejandro Cartas , Jordi Luque , Petia Radeva , Carlos Segura , Mariella Dimiccoli