English
Related papers

Related papers: R+X: Retrieval and Execution from Everyday Human V…

200 papers

Generalist robots should be able to understand and follow user instructions, but current vision-language-action (VLA) models struggle with following fine-grained commands despite providing a powerful architecture for mapping open-vocabulary…

Robotics · Computer Science 2025-08-20 Catherine Glossop , William Chen , Arjun Bhorkar , Dhruv Shah , Sergey Levine

Robust robotic task execution hinges on the reliable detection of execution failures in order to trigger safe operation modes, recovery strategies, or task replanning. However, many failure detection methods struggle to provide meaningful…

Robotics · Computer Science 2025-09-24 Santosh Thoduka , Sebastian Houben , Juergen Gall , Paul G. Plöger

The generation of humanoid animation from text prompts can profoundly impact animation production and AR/VR experiences. However, existing methods only generate body motion data, excluding facial expressions and hand movements. This…

Computer Vision and Pattern Recognition · Computer Science 2024-09-23 Mingdian Liu , Yilin Liu , Gurunandan Krishnan , Karl S Bayer , Bing Zhou

Zero-shot execution of unseen robotic tasks is important to allowing robots to perform a wide variety of tasks in human environments, but collecting the amounts of data necessary to train end-to-end policies in the real-world is often…

Robotics · Computer Science 2021-07-15 Shohin Mukherjee , Chris Paxton , Arsalan Mousavian , Adam Fishman , Maxim Likhachev , Dieter Fox

Recent work on deep reinforcement learning (DRL) has pointed out that algorithmic information about good policies can be extracted from offline data which lack explicit information about executed actions. For example, videos of humans or…

Computer Vision and Pattern Recognition · Computer Science 2023-08-16 Haozhe Liu , Mingchen Zhuge , Bing Li , Yuhui Wang , Francesco Faccio , Bernard Ghanem , Jürgen Schmidhuber

Learning transferable knowledge from unlabeled video data and applying it in new environments is a fundamental capability of intelligent agents. This work presents VideoWorld 2, which extends VideoWorld and offers the first investigation…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Zhongwei Ren , Yunchao Wei , Xiao Yu , Guixun Luo , Yao Zhao , Bingyi Kang , Jiashi Feng , Xiaojie Jin

Enabling robots to learn novel visuomotor skills in a data-efficient manner remains an unsolved problem with myriad challenges. A popular paradigm for tackling this problem is through leveraging large unlabeled datasets that have many…

Robotics · Computer Science 2023-05-16 Maximilian Du , Suraj Nair , Dorsa Sadigh , Chelsea Finn

The ability to autonomously learn behaviors via direct interactions in uninstrumented environments can lead to generalist robots capable of enhancing productivity or providing care in unstructured settings like homes. Such uninstrumented…

Robotics · Computer Science 2021-11-15 Rutav Shah , Vikash Kumar

When an autonomous robot learns how to execute actions, it is of interest to know if and when the execution policy can be generalised to variations of the learning scenarios. This can inform the robot about the necessity of additional…

Robotics · Computer Science 2021-07-21 Alex Mitrevski , Paul G. Plöger , Gerhard Lakemeyer

Hierarchical policies that combine language and low-level control have been shown to perform impressively long-horizon robotic tasks, by leveraging either zero-shot high-level planners like pretrained language and vision-language models…

We introduce PhysWorld, a framework that enables robot learning from video generation through physical world modeling. Recent video generation models can synthesize photorealistic visual demonstrations from language commands and images,…

In recent years, much progress has been made in learning robotic manipulation policies that follow natural language instructions. Such methods typically learn from corpora of robot-language data that was either collected with specific tasks…

Deep reinforcement learning (RL) algorithms can learn complex robotic skills from raw sensory inputs, but have yet to achieve the kind of broad generalization and applicability demonstrated by deep learning methods in supervised domains. We…

Robotics · Computer Science 2018-12-04 Frederik Ebert , Chelsea Finn , Sudeep Dasari , Annie Xie , Alex Lee , Sergey Levine

Pre-training on Internet data has proven to be a key ingredient for broad generalization in many modern ML systems. What would it take to enable such capabilities in robotic reinforcement learning (RL)? Offline RL methods, which learn from…

Synthetic data generated by video generative models has shown promise for robot learning as a scalable pipeline, but it often suffers from inconsistent action quality due to imperfectly generated videos. Recently, vision-language models…

Robotics · Computer Science 2026-02-24 Seungku Kim , Suhyeok Jang , Byungjun Yoon , Dongyoung Kim , John Won , Jinwoo Shin

Negation is a common linguistic skill that allows human to express what we do NOT want. Naturally, one might expect video retrieval to support natural-language queries with negation, e.g., finding shots of kids sitting on the floor and not…

Multimedia · Computer Science 2022-07-14 Ziyue Wang , Aozhu Chen , Fan Hu , Xirong Li

We present SLoMo: a first-of-its-kind framework for transferring skilled motions from casually captured "in the wild" video footage of humans and animals to legged robots. SLoMo works in three stages: 1) synthesize a physically plausible…

Robotics · Computer Science 2023-09-06 John Z. Zhang , Shuo Yang , Gengshan Yang , Arun L. Bishop , Deva Ramanan , Zachary Manchester

Pretrained language models have significantly improved the performance of downstream language understanding tasks, including extractive question answering, by providing high-quality contextualized word embeddings. However, training question…

Computation and Language · Computer Science 2022-06-29 Hongyin Luo , Shang-Wen Li , Mingye Gao , Seunghak Yu , James Glass

Recent progress in imitation learning from human demonstrations has shown promising results in teaching robots manipulation skills. To further scale up training datasets, recent works start to use portable data collection devices without…

Robotics · Computer Science 2024-10-14 Sirui Chen , Chen Wang , Kaden Nguyen , Li Fei-Fei , C. Karen Liu

Household robots operate in the same space for years. Such robots incrementally build dynamic maps that can be used for tasks requiring remote object localization. However, benchmarks in robot learning often test generalization through…

Robotics · Computer Science 2023-01-31 Gunnar A. Sigurdsson , Jesse Thomason , Gaurav S. Sukhatme , Robinson Piramuthu