中文
相关论文

相关论文: Fine-Grained Egocentric Hand-Object Segmentation: …

200 篇论文

We present a video summarization approach for egocentric or "wearable" camera data. Given hours of video, the proposed method produces a compact storyboard summary of the camera wearer's day. In contrast to traditional keyframe selection…

计算机视觉与模式识别 · 计算机科学 2015-05-20 Yong Jae Lee , Kristen Grauman

We present SEMBED, an approach for embedding an egocentric object interaction video in a semantic-visual graph to estimate the probability distribution over its potential semantic labels. When object interactions are annotated using…

计算机视觉与模式识别 · 计算机科学 2016-08-01 Michael Wray , Davide Moltisanti , Walterio Mayol-Cuevas , Dima Damen

Large-scale egocentric video datasets capture diverse human activities across a wide range of scenarios, offering rich and detailed insights into how humans interact with objects, especially those that require fine-grained dexterous…

ENIGMA-51 is a new egocentric dataset acquired in an industrial scenario by 19 subjects who followed instructions to complete the repair of electrical boards using industrial tools (e.g., electric screwdriver) and equipments (e.g.,…

计算机视觉与模式识别 · 计算机科学 2023-11-28 Francesco Ragusa , Rosario Leonardi , Michele Mazzamuto , Claudia Bonanno , Rosario Scavo , Antonino Furnari , Giovanni Maria Farinella

This paper contributes a new high-quality dataset for hand gesture recognition in hand hygiene systems, named "MFH". Generally, current datasets are not focused on: (i) fine-grained actions; and (ii) data mismatch between different…

计算机视觉与模式识别 · 计算机科学 2021-09-08 Huy Q. Vo , Tuong Do , Vi C. Pham , Duy Nguyen , An T. Duong , Quang D. Tran

Background: Egocentric video has recently emerged as a potential solution for monitoring hand function in individuals living with tetraplegia in the community, especially for its ability to detect functional use in the home environment.…

图像与视频处理 · 电气工程与系统科学 2023-11-22 Andrea Bandini , Mehdy Dousty , Sander L. Hitzig , B. Catharine Craven , Sukhvinder Kalsi-Ryan , José Zariffa

Video-Language Pretraining (VLP), which aims to learn transferable representation to advance a wide range of video-text downstream tasks, has recently received increasing attention. Best performing works rely on large-scale, 3rd-person…

Long-form egocentric video understanding provides rich contextual information and unique insights into long-term human behaviors, holding significant potential for applications in embodied intelligence, long-term activity analysis, and…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Wenqi Zhou , Kai Cao , Hao Zheng , Yunze Liu , Xinyi Zheng , Miao Liu , Per Ola Kristensson , Walterio Mayol-Cuevas , Fan Zhang , Weizhe Lin , Junxiao Shen

We present EgoAllo, a system for human motion estimation from a head-mounted device. Using only egocentric SLAM poses and images, EgoAllo guides sampling from a conditional diffusion model to estimate 3D body pose, height, and hand…

计算机视觉与模式识别 · 计算机科学 2024-12-18 Brent Yi , Vickie Ye , Maya Zheng , Yunqi Li , Lea Müller , Georgios Pavlakos , Yi Ma , Jitendra Malik , Angjoo Kanazawa

Generating videos in the first-person perspective has broad application prospects in the field of augmented reality and embodied intelligence. In this work, we explore the cross-view video prediction task, where given an exo-centric video,…

计算机视觉与模式识别 · 计算机科学 2025-04-17 Jilan Xu , Yifei Huang , Baoqi Pei , Junlin Hou , Qingqiu Li , Guo Chen , Yuejie Zhang , Rui Feng , Weidi Xie

Multimodal Large Language Models (MLLMs) have demonstrated remarkable video reasoning capabilities across diverse tasks. However, their ability to understand human intent at a fine-grained level in egocentric videos remains largely…

计算机视觉与模式识别 · 计算机科学 2026-03-13 Ye Pan , Chi Kit Wong , Yuanhuiyi Lyu , Hanqian Li , Jiahao Huo , Jiacheng Chen , Lutao Jiang , Xu Zheng , Xuming Hu

Robot learning increasingly depends on large and diverse data, yet robot data collection remains expensive and difficult to scale. Egocentric human data offer a promising alternative by capturing rich manipulation behavior across everyday…

In this paper, we address the challenge of understanding human activities from an egocentric perspective. Traditional activity recognition techniques face unique challenges in egocentric videos due to the highly dynamic nature of the head…

计算机视觉与模式识别 · 计算机科学 2025-06-03 Zachary Chavis , Stephen J. Guy , Hyun Soo Park

We present an approach for identifying picturesque highlights from large amounts of egocentric video data. Given a set of egocentric videos captured over the course of a vacation, our method analyzes the videos and looks for images that…

计算机视觉与模式识别 · 计算机科学 2016-01-19 Vinay Bettadapura , Daniel Castro , Irfan Essa

Egocentric manipulation videos are highly challenging due to severe occlusions during interactions and frequent object entries and exits from the camera view as the person moves. Current methods typically focus on recovering either hand or…

计算机视觉与模式识别 · 计算机科学 2026-02-26 Yufei Ye , Jiaman Li , Ryan Rong , C. Karen Liu

Although First Person Vision systems can sense the environment from the user's perspective, they are generally unable to predict his intentions and goals. Since human activities can be decomposed in terms of atomic actions and interactions…

计算机视觉与模式识别 · 计算机科学 2019-04-11 Antonino Furnari , Sebastiano Battiato , Kristen Grauman , Giovanni Maria Farinella

The emergence of low-cost personal mobiles devices and wearable cameras and the increasing storage capacity of video-sharing websites have pushed forward a growing interest towards first-person videos. Since most of the recorded videos…

To enable a safe and effective human-robot cooperation, it is crucial to develop models for the identification of human activities. Egocentric vision seems to be a viable solution to solve this problem, and therefore many works provide deep…

计算机视觉与模式识别 · 计算机科学 2023-03-13 Gabriele Goletto , Mirco Planamente , Barbara Caputo , Giuseppe Averta

Research on egocentric tasks in computer vision has mostly focused on head-mounted cameras, such as fisheye cameras or embedded cameras inside immersive headsets. We argue that the increasing miniaturization of optical sensors will lead to…

计算机视觉与模式识别 · 计算机科学 2025-02-26 Dominik Hollidt , Paul Streli , Jiaxi Jiang , Yasaman Haghighi , Changlin Qian , Xintong Liu , Christian Holz

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Boshen Xu , Yuting Mei , Xinbi Liu , Sipeng Zheng , Qin Jin