English
Related papers

Related papers: Multi-modal Egocentric Activity Recognition using …

200 papers

Recent Multimodal Large Language Models (MLLMs) have shown high potential for spatial reasoning within 3D scenes. However, they typically rely on computationally expensive 3D representations like point clouds or reconstructed Bird's-Eye…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Shuyao Shi , Kang G. Shin

Action recognition from an egocentric viewpoint is a crucial perception task in robotics and enables a wide range of human-robot interactions. While most computer vision approaches prioritize the RGB camera, the Depth modality - which can…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Jyoti Kini , Sarah Fleischer , Ishan Dave , Mubarak Shah

Humans develop visual intelligence through perceiving and interacting with their environment - a self-supervised learning process grounded in egocentric experience. Inspired by this, we ask how can artificial systems learn stable object…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yuting Tan , Xilong Cheng , Yunxiao Qin , Zhengnan Li , Jingjing Zhang

Recently, self-supervised pre-training has shown significant improvements in many areas of machine learning, including speech and NLP. We propose using large self-supervised pre-trained models for both audio and text modality with…

Audio and Speech Processing · Electrical Eng. & Systems 2021-08-24 Krishna D N

Automatic personality recognition is a research hotspot in the intersection of computer science and psychology, and in human-computer interaction, personalised has a wide range of applications services and other scenarios. In this paper, an…

Sound · Computer Science 2025-03-11 Weixuan Kong , Jinpeng Yu , Zijun Li , Hanwei Liu , Jiqing Qu , Hui Xiao , Xuefeng Li

Egocentric action recognition is essential for healthcare and assistive technology that relies on egocentric cameras because it allows for the automatic and continuous monitoring of activities of daily living (ADLs) without requiring any…

Computer Vision and Pattern Recognition · Computer Science 2023-06-09 Wiktor Mucha , Martin Kampel

Humans possess a remarkable ability to integrate auditory and visual information, enabling a deeper understanding of the surrounding environment. This early fusion of audio and visual cues, demonstrated through cognitive psychology and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-05 Shentong Mo , Pedro Morgado

In this paper, we propose a novel approach to enhance the 3D body pose estimation of a person computed from videos captured from a single wearable camera. The key idea is to leverage high-level features linking first- and third-views in a…

Computer Vision and Pattern Recognition · Computer Science 2022-06-16 Ameya Dhamanaskar , Mariella Dimiccoli , Enric Corona , Albert Pumarola , Francesc Moreno-Noguer

Recognizing human activities in videos is challenging due to the spatio-temporal complexity and context-dependence of human interactions. Prior studies often rely on single input modalities, such as RGB or skeletal data, limiting their…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Tuyen Tran , Thao Minh Le , Hung Tran , Truyen Tran

As the demand for analyzing egocentric videos grows, egocentric visual attention prediction, anticipating where a camera wearer will attend, has garnered increasing attention. However, it remains challenging due to the inherent complexity…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Sungjune Park , Hongda Mao , Qingshuang Chen , Yong Man Ro , Yelin Kim

Event cameras are neuromorphically inspired sensors that sparsely and asynchronously report brightness changes. Their unique characteristics of high temporal resolution, high dynamic range, and low power consumption make them well-suited…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Haitao Meng , Chonghao Zhong , Sheng Tang , Lian JunJia , Wenwei Lin , Zhenshan Bing , Yi Chang , Gang Chen , Alois Knoll

Egocentric cameras are becoming increasingly popular and provide us with large amounts of videos, captured from the first person perspective. At the same time, surveillance cameras and drones offer an abundance of visual information, often…

Computer Vision and Pattern Recognition · Computer Science 2016-08-16 Shervin Ardeshir , Ali Borji

We introduce an object-aware decoder for improving the performance of spatio-temporal representations on ego-centric videos. The key idea is to enhance object-awareness during training by tasking the model to predict hand positions, object…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Chuhan Zhang , Ankush Gupta , Andrew Zisserman

We present GLNet, a self-supervised framework for learning depth, optical flow, camera pose and intrinsic parameters from monocular video - addressing the difficulty of acquiring realistic ground-truth for such tasks. We propose three…

Computer Vision and Pattern Recognition · Computer Science 2019-09-10 Yuhua Chen , Cordelia Schmid , Cristian Sminchisescu

We address the problem of text-based activity retrieval in video. Given a sentence describing an activity, our task is to retrieve matching clips from an untrimmed video. To capture the inherent structures present in both text and video, we…

Computer Vision and Pattern Recognition · Computer Science 2018-12-27 Huijuan Xu , Kun He , Bryan A. Plummer , Leonid Sigal , Stan Sclaroff , Kate Saenko

Long context egocentric video understanding has recently attracted significant research attention, with augmented reality (AR) highlighted as one of its most important application domains. Nevertheless, the task remains highly challenging…

Machine Learning · Computer Science 2026-04-10 Qiance Tang , Ziqi Wang , Jieyu Lin , Ziyun Li , Barbara De Salvo , Sai Qian Zhang

Self-supervised monocular depth estimation has emerged as a promising approach since it does not rely on labeled training data. Most methods combine convolution and Transformer to model long-distance dependencies to estimate depth…

Computer Vision and Pattern Recognition · Computer Science 2024-09-27 Xuezhi Xiang , Yao Wang , Lei Zhang , Denis Ombati , Himaloy Himu , Xiantong Zhen

We propose a novel system for active semi-supervised feature-based action recognition. Given time sequences of features tracked during movements our system clusters the sequences into actions. Our system is based on encoder-decoder…

Computer Vision and Pattern Recognition · Computer Science 2020-06-15 Jingyuan Li , Eli Shlizerman

Ego-centric data streams provide a unique opportunity to reason about joint behavior by pooling data across individuals. This is especially evident in urban environments teeming with human activities, but which suffer from incomplete and…

Computer Vision and Pattern Recognition · Computer Science 2017-09-06 Guy Rosman , John W. Fisher , Daniela Rus

Learning to infer labels in an open world, i.e., in an environment where the target "labels" are unknown, is an important characteristic for achieving autonomy. Foundation models pre-trained on enormous amounts of data have shown remarkable…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Sanjoy Kundu , Shubham Trehan , Sathyanarayanan N. Aakur