中文
相关论文

相关论文: LOME: Learning Human-Object Manipulation with Acti…

200 篇论文

Loco-manipulation is a fundamental challenge for humanoid robots to achieve versatile interactions in human environments. Although recent studies have made significant progress in humanoid whole-body control, loco-manipulation remains…

机器人学 · 计算机科学 2025-10-14 Yuhui Fu , Feiyang Xie , Chaoyi Xu , Jing Xiong , Haoqi Yuan , Zongqing Lu

We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for…

机器人学 · 计算机科学 2024-10-16 Jinhan Li , Yifeng Zhu , Yuqi Xie , Zhenyu Jiang , Mingyo Seo , Georgios Pavlakos , Yuke Zhu

Extended reality (XR) demands generative models that respond to users' tracked real-world motion, yet current video world models accept only coarse control signals such as text or keyboard input, limiting their utility for embodied…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Linxi Xie , Lisong C. Sun , Ashley Neall , Tong Wu , Shengqu Cai , Gordon Wetzstein

This work presents an object-centric approach to learning vision-based manipulation skills from human videos. We investigate the problem of robot manipulation via imitation in the open-world setting, where a robot learns to manipulate novel…

机器人学 · 计算机科学 2025-09-05 Yifeng Zhu , Arisrei Lim , Peter Stone , Yuke Zhu

Manipulation tasks in daily life, such as pouring water, unfold intentionally under specialized manipulation contexts. Being able to process contextual knowledge in these Activities of Daily Living (ADLs) over time can help us understand…

计算机视觉与模式识别 · 计算机科学 2020-03-04 Chen Jiang , Masood Dehghan , Martin Jagersand

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Jing Tan , Zhaoyang Zhang , Yantao Shen , Jiarui Cai , Shuo Yang , Jiajun Wu , Wei Xia , Zhuowen Tu , Stefano Soatto

Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot datasets, current state-of-the-art video models still…

Large Language Models (LLMs) are gaining popularity in the field of robotics. However, LLM-based robots are limited to simple, repetitive motions due to the poor integration between language models, robots, and the environment. This paper…

机器人学 · 计算机科学 2024-07-02 Haokun Liu , Yaonan Zhu , Kenji Kato , Atsushi Tsukahara , Izumi Kondo , Tadayoshi Aoyama , Yasuhisa Hasegawa

Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Ilya A. Petrov , Vladimir Guzov , Riccardo Marin , Emre Aksan , Xu Chen , Daniel Cremers , Thabo Beeler , Gerard Pons-Moll

Robotic generalization relies on physical intelligence: the ability to reason about state changes, contact-rich interactions, and long-horizon planning under egocentric perception and action. Vision Language Models (VLMs) are essential to…

Humans learn from observations and experiences to adjust their behaviours towards better performance. Interacting with such dynamic humans is challenging, as the robot needs to predict the humans accurately for safe and efficient…

机器人学 · 计算机科学 2025-02-13 Yuwen Liao , Muqing Cao , Xinhang Xu , Lihua Xie

Action-conditioned world models (ACWMs) have shown strong promise for video prediction and decision-making. However, existing benchmarks are largely restricted to egocentric navigation or narrow, task-specific robotics datasets, offering…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Haotian Xue , Yipu Chen , Liqian Ma , Zelin Zhao , Lama Moukheiber , Yuchen Zhu , Yongxin Chen

Loco-manipulation, physical interaction of various objects that is concurrently coordinated with locomotion, remains a major challenge for legged robots due to the need for both precise end-effector control and robustness to unmodeled…

机器人学 · 计算机科学 2025-08-07 Jin Cheng , Dongho Kang , Gabriele Fadini , Guanya Shi , Stelian Coros

Optimizing behaviors for dexterous manipulation has been a longstanding challenge in robotics, with a variety of methods from model-based control to model-free reinforcement learning having been previously explored in literature. Perhaps…

机器人学 · 计算机科学 2022-03-25 Sridhar Pandian Arunachalam , Sneha Silwal , Ben Evans , Lerrel Pinto

Generating realistic human motion is essential for many computer vision and graphics applications. The wide variety of human body shapes and sizes greatly impacts how people move. However, most existing motion models ignore these…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Shashank Tripathi , Omid Taheri , Christoph Lassner , Michael J. Black , Daniel Holden , Carsten Stoll

Human actions involving hand manipulations are structured according to the making and breaking of hand-object contact, and human visual understanding of action is reliant on anticipation of contact as is demonstrated by pioneering work in…

计算机视觉与模式识别 · 计算机科学 2021-02-02 Eadom Dessalene , Chinmaya Devaraj , Michael Maynord , Cornelia Fermuller , Yiannis Aloimonos

Learning an accurate model of the environment is essential for model-based control tasks. Existing methods in robotic visuomotor control usually learn from data with heavily labelled actions, object entities or locations, which can be…

机器人学 · 计算机科学 2021-07-27 Haoqi Yuan , Ruihai Wu , Andrew Zhao , Haipeng Zhang , Zihan Ding , Hao Dong

We present DOME, a novel method for one-shot imitation learning, where a task can be learned from just a single demonstration and then be deployed immediately, without any further data collection or training. DOME does not require prior…

机器人学 · 计算机科学 2022-07-29 Eugene Valassakis , Georgios Papagiannis , Norman Di Palo , Edward Johns

We address the challenging task of detecting the precise moment when hands make contact with objects in egocentric videos. This frame-level detection is crucial for augmented reality, human-computer interaction, assistive technologies, and…

计算机视觉与模式识别 · 计算机科学 2026-04-15 Huy Anh Nguyen , Feras Dayoub , Minh Hoai

Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio…

计算机视觉与模式识别 · 计算机科学 2024-07-26 Changan Chen , Puyuan Peng , Ami Baid , Zihui Xue , Wei-Ning Hsu , David Harwath , Kristen Grauman