中文
相关论文

相关论文: Negative Frames Matter in Egocentric Visual Query …

200 篇论文

Tracking 3D human motion from egocentric multi-camera headset is challenged by severe egomotion, partial visibility or occlusions and lack of training data. Existing methods designed for monocular video often require static or slowly-moving…

计算机视觉与模式识别 · 计算机科学 2026-05-08 Nan Yang , Julian Straub , Fan Zhang , Richard Newcombe , Jakob Engel , Lingni Ma

Augmented reality (AR) games, particularly those designed for head-mounted displays, have grown increasingly prevalent. However, most existing systems depend on pre-scanned, static environments and rely heavily on continuous tracking or…

人机交互 · 计算机科学 2026-02-06 Liuchuan Yu , Ching-I Huang , Hsueh-Cheng Wang , Lap-Fai Yu

Egocentric video provides a unique view into human perception and interaction, with growing relevance for augmented reality, robotics, and assistive technologies. However, rapid camera motion and complex scene dynamics pose major challenges…

计算机视觉与模式识别 · 计算机科学 2026-04-28 Jan Warchocki , Xi Wang , Jonas Kulhanek , Jan van Gemert

Understanding multimodal signals in egocentric vision, such as RGB video, depth, camera poses, and gaze, is essential for applications in augmented reality, robotics, and human-computer interaction, enabling systems to better interpret the…

计算机视觉与模式识别 · 计算机科学 2025-07-22 Gen Li , Yutong Chen , Yiqian Wu , Kaifeng Zhao , Marc Pollefeys , Siyu Tang

This paper investigates the problem of understanding dynamic 3D scenes from egocentric observations, a key challenge in robotics and embodied AI. Unlike prior studies that explored this as long-form video understanding and utilized…

计算机视觉与模式识别 · 计算机科学 2025-01-10 Yue Fan , Xiaojian Ma , Rongpeng Su , Jun Guo , Rujie Wu , Xi Chen , Qing Li

We propose a novel unsupervised approach based on a two-stage object-centric adversarial framework that only needs object regions for detecting frame-level local anomalies in videos. The first stage consists in learning the correspondence…

计算机视觉与模式识别 · 计算机科学 2020-11-16 Pankaj Raj Roy , Guillaume-Alexandre Bilodeau , Lama Seoud

The estimation of optical flow and 6-DoF ego-motion, two fundamental tasks in 3D vision, has typically been addressed independently. For neuromorphic vision (e.g., event cameras), however, the lack of robust data association makes solving…

计算机视觉与模式识别 · 计算机科学 2025-10-27 Wenpu Li , Bangyan Liao , Yi Zhou , Qi Xu , Pian Wan , Peidong Liu

Capturing the state changes of interacting objects is a key technology for understanding human-object interactions. This technical report describes our method using heterogeneous backbones for the Ego4D Object State Change Classification…

计算机视觉与模式识别 · 计算机科学 2022-11-17 Yin-Dong Zheng , Guo Chen , Jiahao Wang , Tong Lu , Limin Wang

Next-generation visual assistants, such as smart glasses, embodied agents, and always-on life-logging systems, must reason over an entire day or more of continuous visual experience. In ultra-long video settings, relevant information is…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Ziyang Wang , Yue Zhang , Shoubin Yu , Ce Zhang , Zengqi Zhao , Jaehong Yoon , Hyunji Lee , Gedas Bertasius , Mohit Bansal

While head-mounted devices are becoming more compact, they provide egocentric views with significant self-occlusions of the device user. Hence, existing methods often fail to accurately estimate complex 3D poses from egocentric views. In…

计算机视觉与模式识别 · 计算机科学 2024-05-16 Hiroyasu Akada , Jian Wang , Vladislav Golyanik , Christian Theobalt

Egocentric video-language understanding demands both high efficiency and accurate spatial-temporal modeling. Existing approaches face three key challenges: 1) Excessive pre-training cost arising from multi-stage pre-training pipelines, 2)…

计算机视觉与模式识别 · 计算机科学 2025-06-18 Xiaoqi Wang , Yi Wang , Lap-Pui Chau

Augmented reality (AR) displays become more and more popular recently, because of its high intuitiveness for humans and high-quality head-mounted display have rapidly developed. To achieve such displays with augmented information, highly…

计算机视觉与模式识别 · 计算机科学 2015-06-22 Kuan-Wen Chen , Chun-Hsin Wang , Xiao Wei , Qiao Liang , Ming-Hsuan Yang , Chu-Song Chen , Yi-Ping Hung

In this paper we propose an end-to-end trainable deep neural network model for egocentric activity recognition. Our model is built on the observation that egocentric activities are highly characterized by the objects and their locations in…

计算机视觉与模式识别 · 计算机科学 2018-08-01 Swathikiran Sudhakaran , Oswald Lanz

Object tracking is an important functionality of edge video analytic systems and services. Multi-object tracking (MOT) detects the moving objects and tracks their locations frame by frame as real scenes are being captured into a video.…

计算机视觉与模式识别 · 计算机科学 2023-09-07 Sanjana Vijay Ganesh , Yanzhao Wu , Gaowen Liu , Ramana Kompella , Ling Liu

We present a visual localization framework based on novel deep attention aware features for autonomous driving that achieves centimeter level localization accuracy. Conventional approaches to the visual localization problem rely on…

计算机视觉与模式识别 · 计算机科学 2020-07-14 Yao Zhou , Guowei Wan , Shenhua Hou , Li Yu , Gang Wang , Xiaofei Rui , Shiyu Song

This paper addresses the daily challenges encountered by visually impaired individuals, such as limited access to information, navigation difficulties, and barriers to social interaction. To alleviate these challenges, we introduce a novel…

计算机视觉与模式识别 · 计算机科学 2024-05-31 Inpyo Song , Minjun Joo , Joonhyung Kwon , Jangwon Lee

Estimating human pose using a front-facing egocentric camera is essential for applications such as sports motion analysis, VR/AR, and AI for wearable devices. However, many existing methods rely on RGB cameras and do not account for…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Wataru Ikeda , Masashi Hatano , Ryosei Hara , Mariko Isogawa

Understanding 3D spatial relationships remains a major limitation of current Vision-Language Models (VLMs). Prior work has addressed this issue by creating spatial question-answering (QA) datasets based on single images or indoor videos.…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Mohsen Gholami , Ahmad Rezaei , Zhou Weimin , Sitong Mao , Shunbo Zhou , Yong Zhang , Mohammad Akbari

Estimating 3D human motion from an egocentric video sequence plays a critical role in human behavior understanding and has various applications in VR/AR. However, naively learning a mapping between egocentric videos and human motions is…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Jiaman Li , C. Karen Liu , Jiajun Wu

Egocentric videos present unique challenges for 3D reconstruction due to rapid camera motion and frequent dynamic interactions. State-of-the-art static reconstruction systems, such as MapAnything, often degrade in these settings, suffering…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Qifei Cui , Patrick Chen