English
Related papers

Related papers: PCIE_LAM Solution for Ego4D Looking At Me Challeng…

200 papers

This report presents our team's PCIE_Interaction solution for the Ego4D Social Interaction Challenge at CVPR 2025, addressing both Looking At Me (LAM) and Talking To Me (TTM) tasks. The challenge requires accurate detection of social…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Kanokphan Lertniphonphan , Feng Chen , Junda Xu , Fengbu Lan , Jun Xie , Tao Zhang , Zhepeng Wang

This report presents our solution to the Ego4D Natural Language Queries (NLQ) Challenge at CVPR 2025. Egocentric video captures the scene from the wearer's perspective, where gaze serves as a key non-verbal communication cue that reflects…

Computer Vision and Pattern Recognition · Computer Science 2025-06-09 Wei-Cheng Lin , Chih-Ming Lien , Chen Lo , Chia-Hung Yeh

In this report, we present our champion solutions to five tracks at Ego4D challenge. We leverage our developed InternVideo, a video foundation model, for five Ego4D tasks, including Moment Queries, Natural Language Queries, Future Hand…

Computer Vision and Pattern Recognition · Computer Science 2022-11-18 Guo Chen , Sen Xing , Zhe Chen , Yi Wang , Kunchang Li , Yizhuo Li , Yi Liu , Jiahao Wang , Yin-Dong Zheng , Bingkun Huang , Zhiyu Zhao , Junting Pan , Yifei Huang , Zun Wang , Jiashuo Yu , Yinan He , Hongjie Zhang , Tong Lu , Yali Wang , Limin Wang , Yu Qiao

The recently released Ego4D dataset and benchmark significantly scales and diversifies the first-person visual perception data. In Ego4D, the Visual Queries 2D Localization task aims to retrieve objects appeared in the past from the…

Computer Vision and Pattern Recognition · Computer Science 2022-08-04 Mengmeng Xu , Cheng-Yang Fu , Yanghao Li , Bernard Ghanem , Juan-Manuel Perez-Rua , Tao Xiang

This report introduces our team's (PCIE_EgoPose) solutions for the EgoExo4D Pose and Proficiency Estimation Challenges at CVPR2025. Focused on the intricate task of estimating 21 3D hand joints from RGB egocentric videos, which are…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Feng Chen , Kanokphan Lertniphonphan , Qiancheng Yan , Xiaohui Fan , Jun Xie , Tao Zhang , Zhepeng Wang

This report presents our team's 'PCIE_EgoHandPose' solution for the EgoExo4D Hand Pose Challenge at CVPR2024. The main goal of the challenge is to accurately estimate hand poses, which involve 21 3D joints, using an RGB egocentric video…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Feng Chen , Ling Ding , Kanokphan Lertniphonphan , Jian Li , Kaer Huang , Zhepeng Wang

In this report, we present the transferring pretrained video mask autoencoders(VideoMAE) to egocentric tasks for Ego4d Looking at me Challenge. VideoMAE is the data-efficient pretraining model for self-supervised video pre-training and can…

Computer Vision and Pattern Recognition · Computer Science 2022-11-30 Yinan He , Guo Chen

Accurately estimating the 3D pose of the camera wearer in egocentric video sequences is crucial to modeling human behavior in virtual and augmented reality applications. The task presents unique challenges due to the limited visibility of…

Computer Vision and Pattern Recognition · Computer Science 2024-11-08 Luca Scofano , Alessio Sampieri , Edoardo De Matteis , Indro Spinelli , Fabio Galasso

In this report, we propose a video-language pretraining (VLP) based solution \cite{kevin2022egovlp} for four Ego4D challenge tasks, including Natural Language Query (NLQ), Moment Query (MQ), Object State Change Classification (OSCC), and…

The ability to anticipate human-object interactions is highly desirable in an intelligent assistive system in order to guide users during daily life activities and understand their short and long-term goals. Creating systems with such…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Daniele Materia , Francesco Ragusa , Giovanni Maria Farinella

An embodied AI assistant operating on egocentric video must integrate spatial cues across time - for instance, determining where an object A, glimpsed a few moments ago lies relative to an object B encountered later. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Sahithya Ravi , Gabriel Sarch , Vibhav Vineet , Andrew D. Wilson , Balasaravanan Thoravi Kumaravel

Facial micro-expressions (MEs) are involuntary movements of the face that occur spontaneously when a person experiences an emotion but attempts to suppress or repress the facial expression, typically found in a high-stakes environment. In…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Xinqi Fan , Jingting Li , John See , Moi Hoon Yap , Wen-Huang Cheng , Xiaobai Li , Xiaopeng Hong , Su-Jing Wang , Adrian K. Davision

In this report, we present our approach for the Natural Language Query track and Goal Step track of the Ego4D Episodic Memory Benchmark at CVPR 2024. Both challenges require the localization of actions within long video sequences using…

Computer Vision and Pattern Recognition · Computer Science 2024-11-19 Yisen Feng , Haoyu Zhang , Yuquan Xie , Zaijing Li , Meng Liu , Liqiang Nie

Gaze and face tracking algorithms have traditionally battled a compromise between computational complexity and accuracy; the most accurate neural net algorithms cannot be implemented in real time, but less complex real-time algorithms…

Computer Vision and Pattern Recognition · Computer Science 2017-11-21 George He , Sami Oueida , Tucker Ward

In this report, we present our approach and empirical results of applying masked autoencoders in two egocentric video understanding tasks, namely, Object State Change Classification and PNR Temporal Localization, of Ego4D Challenge 2022. As…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Jiachen Lei , Shuang Ma , Zhongjie Ba , Sai Vemprala , Ashish Kapoor , Kui Ren

Understanding human activity is a crucial yet intricate task in egocentric vision, a field that focuses on capturing visual perspectives from the camera wearer's viewpoint. Traditional methods heavily rely on representation learning that is…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Sanghwan Kim , Daoji Huang , Yongqin Xian , Otmar Hilliges , Luc Van Gool , Xi Wang

Existing solutions to visual simultaneous localization and mapping (V-SLAM) assume that errors in feature extraction and matching are independent and identically distributed (i.i.d), but this assumption is known to not be true -- features…

Computer Vision and Pattern Recognition · Computer Science 2020-11-20 Sadegh Rabiee , Joydeep Biswas

Visual queries 3D localization (VQ3D) is a task in the Ego4D Episodic Memory Benchmark. Given an egocentric video, the goal is to answer queries of the form "Where did I last see object X?", where the query object X is specified as a static…

Computer Vision and Pattern Recognition · Computer Science 2022-11-21 Jinjie Mai , Chen Zhao , Abdullah Hamdi , Silvio Giancola , Bernard Ghanem

Generating captions for long and complex videos is both critical and challenging, with significant implications for the growing fields of text-to-video generation and multi-modal understanding. One key challenge in long video captioning is…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Zhantao Yang , Huangji Wang , Ruili Feng , Han Zhang , Yuting Hu , Shangwen Zhu , Junyan Li , Yu Liu , Fan Cheng

Lifelogging devices are spreading faster everyday. This growth can represent great benefits to develop methods for extraction of meaningful information about the user wearing the device and his/her environment. In this paper, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2015-07-09 Marc Bolaños , Petia Radeva
‹ Prev 1 2 3 10 Next ›