中文
相关论文

相关论文: EgoAVFlow: Robot Policy Learning with Active Visio…

200 篇论文

Real robot data collection for imitation learning has led to significant advancements in robotic manipulation. However, the requirement for robot hardware in the process fundamentally constrains the scale of the data. In this paper, we…

Egocentric videos capture how humans manipulate objects and tools, providing diverse motion cues for learning object manipulation. Unlike the costly, expert-driven manual teleoperation commonly used in training Vision-Language-Action models…

机器人学 · 计算机科学 2025-09-29 Tomoya Yoshida , Shuhei Kurita , Taichi Nishimura , Shinsuke Mori

Imitation learning based visuomotor policies have achieved strong performance in robotic manipulation, yet they often remain sensitive to egocentric viewpoint shifts. Unlike third-person viewpoint changes that only move the camera,…

We present an approach to robot learning from egocentric human videos by modeling human preferences in a reward function and optimizing robot behavior to maximize this reward. Prior work on reward learning from human videos attempts to…

机器人学 · 计算机科学 2026-02-13 Mrinal Verghese , Christopher G. Atkeson

Robots operating in complex and uncertain environments face considerable challenges. Advanced robotic systems often rely on extensive datasets to learn manipulation tasks. In contrast, when humans are faced with unfamiliar tasks, such as…

机器人学 · 计算机科学 2025-11-10 Yichen Zhu , Feifei Feng

First-person video highlights a camera-wearer's activities in the context of their persistent environment. However, current video understanding approaches reason over visual features from short video clips that are detached from the…

计算机视觉与模式识别 · 计算机科学 2023-11-13 Tushar Nagarajan , Santhosh Kumar Ramakrishnan , Ruta Desai , James Hillis , Kristen Grauman

Egocentric video-language pretraining has significantly advanced video representation learning. Humans perceive and interact with a fully 3D world, developing spatial awareness that extends beyond text-based understanding. However, most…

计算机视觉与模式识别 · 计算机科学 2025-12-05 Boshen Xu , Yuting Mei , Xinbi Liu , Sipeng Zheng , Qin Jin

Egocentric video generation with fine-grained control through body motion is a key requirement towards embodied AI agents that can simulate, predict, and plan actions. In this work, we propose EgoControl, a pose-controllable video diffusion…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Enrico Pallotta , Sina Mokhtarzadeh Azar , Lars Doorenbos , Serdar Ozsoy , Umar Iqbal , Juergen Gall

Emerging embodied AI applications, such as wearable cameras and autonomous agents, have underscored the need for robust reasoning from first person video streams. We introduce EgoVLM, a vision-language model specifically designed to…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Ashwin Vinod , Shrey Pandit , Aditya Vavre , Linshen Liu

Understanding and predicting object motion from egocentric video is fundamental to embodied perception and interaction. However, generating physically consistent 6DoF trajectories remains challenging due to occlusions, fast motion, and the…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Abhishek Saroha , Huajian Zeng , Xingxing Zuo , Daniel Cremers , Xi Wang

Egocentric videos can bring a lot of information about how humans perceive the world and interact with the environment, which can be beneficial for the analysis of human behaviour. The research in egocentric video analysis is developing…

计算机视觉与模式识别 · 计算机科学 2021-07-29 Ivan Rodin , Antonino Furnari , Dimitrios Mavroedis , Giovanni Maria Farinella

Human demonstrations offer rich environmental diversity and scale naturally, making them an appealing alternative to robot teleoperation. While this paradigm has advanced robot-arm manipulation, its potential for the more challenging,…

机器人学 · 计算机科学 2026-02-11 Modi Shi , Shijia Peng , Jin Chen , Haoran Jiang , Yinghui Li , Di Huang , Ping Luo , Hongyang Li , Li Chen

Wearable collaborative robots stand to assist human wearers who need fall prevention assistance or wear exoskeletons. Such a robot needs to be able to constantly adapt to the surrounding scene based on egocentric vision, and predict the ego…

计算机视觉与模式识别 · 计算机科学 2024-08-08 Weizhuo Wang , C. Karen Liu , Monroe Kennedy

Imitation learning from human demonstrations offers a promising approach for robot skill acquisition, but egocentric human data introduces fundamental challenges due to the embodiment gap. During manipulation, humans actively coordinate…

机器人学 · 计算机科学 2026-03-11 Justin Yu , Yide Shentu , Di Wu , Pieter Abbeel , Ken Goldberg , Philipp Wu

The advancement of robot learning is currently hindered by the scarcity of large-scale, high-quality datasets. While established data collection methods such as teleoperation and universal manipulation interfaces dominate current datasets,…

We present EgoNav, a system that enables a humanoid robot to traverse diverse, unseen environments by learning entirely from 5 hours of human walking data, with no robot data or finetuning. A diffusion model predicts distributions of…

机器人学 · 计算机科学 2026-04-02 Weizhuo Wang , Yanjie Ze , C. Karen Liu , Monroe Kennedy

Robust and accurate perception of humans in their 3D scene context is essential for integrating robots into everyday environments. Existing approaches, however, often fail to predict plausible and accurate human motion estimates that are…

机器人学 · 计算机科学 2026-05-26 Simon Schaefer , Joshua Näf , Stefan Leutenegger

Learning robot control policies from human videos is a promising direction for scaling up robot learning. However, how to extract action knowledge (or action representations) from videos for policy learning remains a key challenge. Existing…

机器人学 · 计算机科学 2025-06-05 Zhao-Heng Yin , Sherry Yang , Pieter Abbeel

The scale and diversity of demonstration data required for imitation learning is a significant challenge. We present EgoMimic, a full-stack framework which scales manipulation via human embodiment data, specifically egocentric human videos…

机器人学 · 计算机科学 2024-11-01 Simar Kareer , Dhruv Patel , Ryan Punamiya , Pranay Mathur , Shuo Cheng , Chen Wang , Judy Hoffman , Danfei Xu

Given a video captured from a first person perspective and the environment context of where the video is recorded, can we recognize what the person is doing and identify where the action occurs in the 3D space? We address this challenging…

计算机视觉与模式识别 · 计算机科学 2022-08-16 Miao Liu , Lingni Ma , Kiran Somasundaram , Yin Li , Kristen Grauman , James M. Rehg , Chao Li
‹ 上一页 1 2 3 10 下一页 ›