English
Related papers

Related papers: LOME: Learning Human-Object Manipulation with Acti…

200 papers

Loco-manipulation is a fundamental challenge for humanoid robots to achieve versatile interactions in human environments. Although recent studies have made significant progress in humanoid whole-body control, loco-manipulation remains…

Robotics · Computer Science 2025-10-14 Yuhui Fu , Feiyang Xie , Chaoyi Xu , Jing Xiong , Haoqi Yuan , Zongqing Lu

We study the problem of teaching humanoid robots manipulation skills by imitating from single video demonstrations. We introduce OKAMI, a method that generates a manipulation plan from a single RGB-D video and derives a policy for…

Robotics · Computer Science 2024-10-16 Jinhan Li , Yifeng Zhu , Yuqi Xie , Zhenyu Jiang , Mingyo Seo , Georgios Pavlakos , Yuke Zhu

Extended reality (XR) demands generative models that respond to users' tracked real-world motion, yet current video world models accept only coarse control signals such as text or keyboard input, limiting their utility for embodied…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Linxi Xie , Lisong C. Sun , Ashley Neall , Tong Wu , Shengqu Cai , Gordon Wetzstein

This work presents an object-centric approach to learning vision-based manipulation skills from human videos. We investigate the problem of robot manipulation via imitation in the open-world setting, where a robot learns to manipulate novel…

Robotics · Computer Science 2025-09-05 Yifeng Zhu , Arisrei Lim , Peter Stone , Yuke Zhu

Manipulation tasks in daily life, such as pouring water, unfold intentionally under specialized manipulation contexts. Being able to process contextual knowledge in these Activities of Daily Living (ADLs) over time can help us understand…

Computer Vision and Pattern Recognition · Computer Science 2020-03-04 Chen Jiang , Masood Dehghan , Martin Jagersand

We introduce Talk2Move, a reinforcement learning (RL) based diffusion framework for text-instructed spatial transformation of objects within scenes. Spatially manipulating objects in a scene through natural language poses a challenge for…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Jing Tan , Zhaoyang Zhang , Yantao Shen , Jiarui Cai , Shuo Yang , Jiajun Wu , Wei Xia , Zhuowen Tu , Stefano Soatto

Action-conditioned video models offer a promising path to building general-purpose robot simulators that can improve directly from data. Yet, despite training on large-scale robot datasets, current state-of-the-art video models still…

Large Language Models (LLMs) are gaining popularity in the field of robotics. However, LLM-based robots are limited to simple, repetitive motions due to the poor integration between language models, robots, and the environment. This paper…

Modeling human-object interactions (HOI) from an egocentric perspective is a critical yet challenging task, particularly when relying on sparse signals from wearable devices like smart glasses and watches. We present ECHO, the first unified…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ilya A. Petrov , Vladimir Guzov , Riccardo Marin , Emre Aksan , Xu Chen , Daniel Cremers , Thabo Beeler , Gerard Pons-Moll

Robotic generalization relies on physical intelligence: the ability to reason about state changes, contact-rich interactions, and long-horizon planning under egocentric perception and action. Vision Language Models (VLMs) are essential to…

Humans learn from observations and experiences to adjust their behaviours towards better performance. Interacting with such dynamic humans is challenging, as the robot needs to predict the humans accurately for safe and efficient…

Robotics · Computer Science 2025-02-13 Yuwen Liao , Muqing Cao , Xinhang Xu , Lihua Xie

Action-conditioned world models (ACWMs) have shown strong promise for video prediction and decision-making. However, existing benchmarks are largely restricted to egocentric navigation or narrow, task-specific robotics datasets, offering…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Haotian Xue , Yipu Chen , Liqian Ma , Zelin Zhao , Lama Moukheiber , Yuchen Zhu , Yongxin Chen

Loco-manipulation, physical interaction of various objects that is concurrently coordinated with locomotion, remains a major challenge for legged robots due to the need for both precise end-effector control and robustness to unmodeled…

Robotics · Computer Science 2025-08-07 Jin Cheng , Dongho Kang , Gabriele Fadini , Guanya Shi , Stelian Coros

Optimizing behaviors for dexterous manipulation has been a longstanding challenge in robotics, with a variety of methods from model-based control to model-free reinforcement learning having been previously explored in literature. Perhaps…

Robotics · Computer Science 2022-03-25 Sridhar Pandian Arunachalam , Sneha Silwal , Ben Evans , Lerrel Pinto

Generating realistic human motion is essential for many computer vision and graphics applications. The wide variety of human body shapes and sizes greatly impacts how people move. However, most existing motion models ignore these…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Shashank Tripathi , Omid Taheri , Christoph Lassner , Michael J. Black , Daniel Holden , Carsten Stoll

Human actions involving hand manipulations are structured according to the making and breaking of hand-object contact, and human visual understanding of action is reliant on anticipation of contact as is demonstrated by pioneering work in…

Computer Vision and Pattern Recognition · Computer Science 2021-02-02 Eadom Dessalene , Chinmaya Devaraj , Michael Maynord , Cornelia Fermuller , Yiannis Aloimonos

Learning an accurate model of the environment is essential for model-based control tasks. Existing methods in robotic visuomotor control usually learn from data with heavily labelled actions, object entities or locations, which can be…

Robotics · Computer Science 2021-07-27 Haoqi Yuan , Ruihai Wu , Andrew Zhao , Haipeng Zhang , Zihan Ding , Hao Dong

We present DOME, a novel method for one-shot imitation learning, where a task can be learned from just a single demonstration and then be deployed immediately, without any further data collection or training. DOME does not require prior…

Robotics · Computer Science 2022-07-29 Eugene Valassakis , Georgios Papagiannis , Norman Di Palo , Edward Johns

We address the challenging task of detecting the precise moment when hands make contact with objects in egocentric videos. This frame-level detection is crucial for augmented reality, human-computer interaction, assistive technologies, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Huy Anh Nguyen , Feras Dayoub , Minh Hoai

Generating realistic audio for human actions is important for many applications, such as creating sound effects for films or virtual reality games. Existing approaches implicitly assume total correspondence between the video and audio…

Computer Vision and Pattern Recognition · Computer Science 2024-07-26 Changan Chen , Puyuan Peng , Ami Baid , Zihui Xue , Wei-Ning Hsu , David Harwath , Kristen Grauman
‹ Prev 1 4 5 6 7 8 10 Next ›