中文
相关论文

相关论文: EgoEdit: Dataset, Real-Time Streaming Model, and B…

200 篇论文

Visual feedback is critical for motor skill acquisition in sports and rehabilitation, and psychological studies show that observing near-perfect versions of one's own performance accelerates learning more effectively than watching expert…

计算机视觉与模式识别 · 计算机科学 2026-04-14 Arjun Somayazulu , Kristen Grauman

We propose to forecast future hand-object interactions given an egocentric video. Instead of predicting action labels or pixels, we directly predict the hand motion trajectory and the future contact points on the next active object (i.e.,…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Shaowei Liu , Subarna Tripathi , Somdeb Majumdar , Xiaolong Wang

The rapid development of Multimodal Large Language Models (MLLMs) has led to growing interest in egocentric video understanding, specifically the ability for MLLMs to recognize fine-grained hand-object interactions, track object state…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Yang Dai , Dian Jiao , Tianwei Lin , Wenqiao Zhang

This technical report introduces our solution, MEEV, proposed to the EgoBody Challenge at ECCV 2022. Captured from head-mounted devices, the dataset consists of human body shape and motion of interacting people. The EgoBody dataset has…

计算机视觉与模式识别 · 计算机科学 2022-10-26 Nicolas Monet , Dongyoon Wee

We study the problem of estimating the body movements of a camera wearer from egocentric videos. Current methods for ego-body pose estimation rely on temporally dense sensor data, such as IMU measurements from spatially sparse body parts…

计算机视觉与模式识别 · 计算机科学 2024-11-07 Seunggeun Chi , Pin-Hao Huang , Enna Sachdeva , Hengbo Ma , Karthik Ramani , Kwonjoon Lee

Immersive virtual reality (VR) applications demand accurate, temporally coherent full-body pose tracking. Recent head-mounted camera-based approaches show promise in egocentric pose estimation, but encounter challenges when applied to VR…

计算机视觉与模式识别 · 计算机科学 2026-02-06 Haojie Cheng , Shaun Jing Heng Ong , Shaoyu Cai , Aiden Tat Yang Koh , Fuxi Ouyang , Eng Tat Khoo

Accurately estimating and forecasting human body pose is important for enhancing the user's sense of immersion in Augmented Reality. Addressing this need, our paper introduces EgoCast, a bimodal method for 3D human pose forecasting using…

计算机视觉与模式识别 · 计算机科学 2024-12-05 Maria Escobar , Juanita Puentes , Cristhian Forigua , Jordi Pont-Tuset , Kevis-Kokitsi Maninis , Pablo Arbelaez

Instruction-based video editing is a natural way to control video content with text, but adapting a video generation model into an editor usually appears data-hungry. At the same time, high-quality video editing data remains scarce. In this…

计算机视觉与模式识别 · 计算机科学 2026-04-13 Zhefan Rao , Bin Zou , Haoxuan Che , Xuanhua He , Chong Hou Choi , Yanheng Li , Rui Liu , Qifeng Chen

We propose EgoGrasp, the first method to reconstruct world-space hand-object interactions (W-HOI) from dynamic egoview videos, supporting open-vocabulary objects. Accurate W-HOI reconstruction is critical for embodied intelligence yet…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Hongming Fu , Wenjia Wang , Xiaozhen Qiao , Rolandos Alexandros Potamias , Taku Komura , Shuo Yang , Zheng Liu , Bo Zhao

Large-scale text-to-video models have shown remarkable abilities, but their direct application in video editing remains challenging due to limited available datasets. Current video editing methods commonly require per-video fine-tuning of…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Zhenghao Zhang , Zuozhuo Dai , Long Qin , Weizhi Wang

In recent years, the thriving development of research related to egocentric videos has provided a unique perspective for the study of conversational interactions, where both visual and audio signals play a crucial role. While most prior…

计算机视觉与模式识别 · 计算机科学 2024-04-04 Wenqi Jia , Miao Liu , Hao Jiang , Ishwarya Ananthabhotla , James M. Rehg , Vamsi Krishna Ithapu , Ruohan Gao

Event boundaries play a crucial role as a pre-processing step for detection, localization, and recognition tasks of human activities in videos. Typically, although their intrinsic subjectiveness, temporal bounds are provided manually as…

计算机视觉与模式识别 · 计算机科学 2018-09-07 Alejandro Cartas , Estefania Talavera , Petia Radeva , Mariella Dimiccoli

Egocentric video grounding is a crucial task for embodied intelligence applications, distinct from exocentric video moment localization. Existing methods primarily focus on the distributional differences between egocentric and exocentric…

计算机视觉与模式识别 · 计算机科学 2025-05-08 Yisen Feng , Haoyu Zhang , Meng Liu , Weili Guan , Liqiang Nie

The ability to predict collision-free future trajectories from egocentric observations is crucial in applications such as humanoid robotics, VR / AR, and assistive navigation. In this work, we introduce the challenging problem of predicting…

计算机视觉与模式识别 · 计算机科学 2025-08-21 Boxiao Pan , Adam W. Harley , C. Karen Liu , Leonidas J. Guibas

We introduce FEEL (Force-Enhanced Egocentric Learning), the first large-scale dataset pairing force measurements gathered from custom piezoresistive gloves with egocentric video. Our gloves enable scalable data collection, and FEEL contains…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Eadom Dessalene , Botao He , Michael Maynord , Yonatan Tussa , Pavan Mantripragada , Yianni Karabati , Nirupam Roy , Yiannis Aloimonos

Recent advancements in Multi-modal Large Language Models (MLLMs) have opened new avenues for applications in Embodied AI. Building on previous work, EgoThink, we introduce VidEgoThink, a comprehensive benchmark for evaluating egocentric…

计算机视觉与模式识别 · 计算机科学 2024-10-16 Sijie Cheng , Kechen Fang , Yangyang Yu , Sicheng Zhou , Bohao Li , Ye Tian , Tingguang Li , Lei Han , Yang Liu

Egocentric videos present unique challenges for 3D scene understanding due to rapid camera motion, frequent object occlusions, and limited object visibility. This paper introduces a novel approach to instance segmentation and tracking in…

计算机视觉与模式识别 · 计算机科学 2024-11-21 Yash Bhalgat , Vadim Tschernezki , Iro Laina , João F. Henriques , Andrea Vedaldi , Andrew Zisserman

To serve as a scalable data source for embodied AI, world models should act as true simulators that infer interaction dynamics strictly from user actions, rather than mere conditional video generators relying on privileged future object…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Dayou Li , Lulin Liu , Bangya Liu , Shijie Zhou , Jiu Feng , Ziqi Lu , Minghui Zheng , Chenyu You , Zhiwen Fan

Observing surgical practice has historically relied on fixed vantage points or recollections, leaving the egocentric visual perspectives that guide clinical decisions undocumented. Fixed-camera video can capture surgical workflows at the…

Recently, we have witnessed great progress in image editing with natural language instructions. Several closed-source models like GPT-Image-1, Seedream, and Google-Nano-Banana have shown highly promising progress. However, the open-source…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Keming Wu , Sicong Jiang , Max Ku , Ping Nie , Minghao Liu , Wenhu Chen