English
Related papers

Related papers: RH20T-P: A Primitive-Level Robotic Dataset Towards…

200 papers

Robot learning holds the promise of learning policies that generalize broadly. However, such generalization requires sufficiently diverse datasets of the task of interest, which can be prohibitively expensive to collect. In other fields,…

Graphical user interface (GUI)-based mobile agents automate digital tasks on mobile devices by interpreting natural-language instructions and interacting with the screen. While recent methods apply reinforcement learning (RL) to train…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Li Gu , Zihuan Jiang , Zhixiang Chi , Huan Liu , Ziqiang Wang , Yuanhao Yu , Glen Berseth , Yang Wang

Human videos contain rich manipulation priors, but using them for robot learning remains difficult because raw observations entangle scene understanding, human motion, and embodiment-specific action. We introduce MoT-HRA, a hierarchical…

Robotics · Computer Science 2026-05-22 Yifan Xie , YuAn Wang , Guangyu Chen , Jinkun Liu , Yu Sun , Wenbo Ding

To tackle long-horizon tasks, recent hierarchical vision-language-action (VLAs) frameworks employ vision-language model (VLM)-based planners to decompose complex manipulation tasks into simpler sub-tasks that low-level visuomotor policies…

Robotics · Computer Science 2025-10-17 Mingxuan Yan , Yuping Wang , Zechun Liu , Jiachen Li

Real-world robotic manipulation tasks remain an elusive challenge, since they involve both fine-grained environment interaction, as well as the ability to plan for long-horizon goals. Although deep reinforcement learning (RL) methods have…

Machine Learning · Computer Science 2023-03-20 Núria Armengol Urpí , Marco Bagatella , Otmar Hilliges , Georg Martius , Stelian Coros

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with…

Robotics · Computer Science 2025-05-15 Embodiment Collaboration , Abby O'Neill , Abdul Rehman , Abhinav Gupta , Abhiram Maddukuri , Abhishek Gupta , Abhishek Padalkar , Abraham Lee , Acorn Pooley , Agrim Gupta , Ajay Mandlekar , Ajinkya Jain , Albert Tung , Alex Bewley , Alex Herzog , Alex Irpan , Alexander Khazatsky , Anant Rai , Anchit Gupta , Andrew Wang , Andrey Kolobov , Anikait Singh , Animesh Garg , Aniruddha Kembhavi , Annie Xie , Anthony Brohan , Antonin Raffin , Archit Sharma , Arefeh Yavary , Arhan Jain , Ashwin Balakrishna , Ayzaan Wahid , Ben Burgess-Limerick , Beomjoon Kim , Bernhard Schölkopf , Blake Wulfe , Brian Ichter , Cewu Lu , Charles Xu , Charlotte Le , Chelsea Finn , Chen Wang , Chenfeng Xu , Cheng Chi , Chenguang Huang , Christine Chan , Christopher Agia , Chuer Pan , Chuyuan Fu , Coline Devin , Danfei Xu , Daniel Morton , Danny Driess , Daphne Chen , Deepak Pathak , Dhruv Shah , Dieter Büchler , Dinesh Jayaraman , Dmitry Kalashnikov , Dorsa Sadigh , Edward Johns , Ethan Foster , Fangchen Liu , Federico Ceola , Fei Xia , Feiyu Zhao , Felipe Vieira Frujeri , Freek Stulp , Gaoyue Zhou , Gaurav S. Sukhatme , Gautam Salhotra , Ge Yan , Gilbert Feng , Giulio Schiavi , Glen Berseth , Gregory Kahn , Guangwen Yang , Guanzhi Wang , Hao Su , Hao-Shu Fang , Haochen Shi , Henghui Bao , Heni Ben Amor , Henrik I Christensen , Hiroki Furuta , Homanga Bharadhwaj , Homer Walke , Hongjie Fang , Huy Ha , Igor Mordatch , Ilija Radosavovic , Isabel Leal , Jacky Liang , Jad Abou-Chakra , Jaehyung Kim , Jaimyn Drake , Jan Peters , Jan Schneider , Jasmine Hsu , Jay Vakil , Jeannette Bohg , Jeffrey Bingham , Jeffrey Wu , Jensen Gao , Jiaheng Hu , Jiajun Wu , Jialin Wu , Jiankai Sun , Jianlan Luo , Jiayuan Gu , Jie Tan , Jihoon Oh , Jimmy Wu , Jingpei Lu , Jingyun Yang , Jitendra Malik , João Silvério , Joey Hejna , Jonathan Booher , Jonathan Tompson , Jonathan Yang , Jordi Salvador , Joseph J. Lim , Junhyek Han , Kaiyuan Wang , Kanishka Rao , Karl Pertsch , Karol Hausman , Keegan Go , Keerthana Gopalakrishnan , Ken Goldberg , Kendra Byrne , Kenneth Oslund , Kento Kawaharazuka , Kevin Black , Kevin Lin , Kevin Zhang , Kiana Ehsani , Kiran Lekkala , Kirsty Ellis , Krishan Rana , Krishnan Srinivasan , Kuan Fang , Kunal Pratap Singh , Kuo-Hao Zeng , Kyle Hatch , Kyle Hsu , Laurent Itti , Lawrence Yunliang Chen , Lerrel Pinto , Li Fei-Fei , Liam Tan , Linxi "Jim" Fan , Lionel Ott , Lisa Lee , Luca Weihs , Magnum Chen , Marion Lepert , Marius Memmel , Masayoshi Tomizuka , Masha Itkina , Mateo Guaman Castro , Max Spero , Maximilian Du , Michael Ahn , Michael C. Yip , Mingtong Zhang , Mingyu Ding , Minho Heo , Mohan Kumar Srirama , Mohit Sharma , Moo Jin Kim , Muhammad Zubair Irshad , Naoaki Kanazawa , Nicklas Hansen , Nicolas Heess , Nikhil J Joshi , Niko Suenderhauf , Ning Liu , Norman Di Palo , Nur Muhammad Mahi Shafiullah , Oier Mees , Oliver Kroemer , Osbert Bastani , Pannag R Sanketi , Patrick "Tree" Miller , Patrick Yin , Paul Wohlhart , Peng Xu , Peter David Fagan , Peter Mitrano , Pierre Sermanet , Pieter Abbeel , Priya Sundaresan , Qiuyu Chen , Quan Vuong , Rafael Rafailov , Ran Tian , Ria Doshi , Roberto Martín-Martín , Rohan Baijal , Rosario Scalise , Rose Hendrix , Roy Lin , Runjia Qian , Ruohan Zhang , Russell Mendonca , Rutav Shah , Ryan Hoque , Ryan Julian , Samuel Bustamante , Sean Kirmani , Sergey Levine , Shan Lin , Sherry Moore , Shikhar Bahl , Shivin Dass , Shubham Sonawani , Shubham Tulsiani , Shuran Song , Sichun Xu , Siddhant Haldar , Siddharth Karamcheti , Simeon Adebola , Simon Guist , Soroush Nasiriany , Stefan Schaal , Stefan Welker , Stephen Tian , Subramanian Ramamoorthy , Sudeep Dasari , Suneel Belkhale , Sungjae Park , Suraj Nair , Suvir Mirchandani , Takayuki Osa , Tanmay Gupta , Tatsuya Harada , Tatsuya Matsushima , Ted Xiao , Thomas Kollar , Tianhe Yu , Tianli Ding , Todor Davchev , Tony Z. Zhao , Travis Armstrong , Trevor Darrell , Trinity Chung , Vidhi Jain , Vikash Kumar , Vincent Vanhoucke , Vitor Guizilini , Wei Zhan , Wenxuan Zhou , Wolfram Burgard , Xi Chen , Xiangyu Chen , Xiaolong Wang , Xinghao Zhu , Xinyang Geng , Xiyuan Liu , Xu Liangwei , Xuanlin Li , Yansong Pang , Yao Lu , Yecheng Jason Ma , Yejin Kim , Yevgen Chebotar , Yifan Zhou , Yifeng Zhu , Yilin Wu , Ying Xu , Yixuan Wang , Yonatan Bisk , Yongqiang Dou , Yoonyoung Cho , Youngwoon Lee , Yuchen Cui , Yue Cao , Yueh-Hua Wu , Yujin Tang , Yuke Zhu , Yunchu Zhang , Yunfan Jiang , Yunshuang Li , Yunzhu Li , Yusuke Iwasawa , Yutaka Matsuo , Zehan Ma , Zhuo Xu , Zichen Jeff Cui , Zichen Zhang , Zipeng Fu , Zipeng Lin

"Code-as-Policy" considers how executable code can complement data-intensive Vision-Language-Action (VLA) methods, yet their effectiveness as autonomous controllers for embodied manipulation remains underexplored. We present CaP-X, an…

This work aims to learn how to perform complex robot manipulation tasks that are composed of several, consecutively executed low-level sub-tasks, given as input a few visual demonstrations of the tasks performed by a person. The sub-tasks…

Robotics · Computer Science 2022-03-09 Junchi Liang , Bowen Wen , Kostas Bekris , Abdeslam Boularias

Video generative models demonstrate great promise in robotics by serving as visual planners or as policy supervisors. When pretrained on internet-scale data, such video models intimately understand alignment with natural language, and can…

Machine Learning · Computer Science 2025-04-23 Calvin Luo , Zilai Zeng , Yilun Du , Chen Sun

Large language models (LLMs) are shown to possess a wealth of actionable knowledge that can be extracted for robot manipulation in the form of reasoning and planning. Despite the progress, most still rely on pre-defined motion primitives to…

Robotics · Computer Science 2023-11-03 Wenlong Huang , Chen Wang , Ruohan Zhang , Yunzhu Li , Jiajun Wu , Li Fei-Fei

This paper focuses on embodied task planning, where an agent acquires visual observations from the environment and executes atomic actions to accomplish a given task. Although recent Vision-Language Models (VLMs) have achieved impressive…

Robotics · Computer Science 2026-04-10 Peiran Xu , Jiaqi Zheng , Yadong Mu

Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLAC, a general process reward model built upon InternVL and…

Fine-tuning vision-language models (VLMs) on robot teleoperation data to create vision-language-action (VLA) models is a promising paradigm for training generalist policies, but it suffers from a fundamental tradeoff: learning to produce…

Robotics · Computer Science 2025-09-29 Asher J. Hancock , Xindi Wu , Lihan Zha , Olga Russakovsky , Anirudha Majumdar

Teaching robots desired skills in real-world environments remains challenging, especially for non-experts. A key bottleneck is that collecting robotic data often requires expertise or specialized hardware, limiting accessibility and…

Robotics · Computer Science 2025-05-13 Gi-Cheon Kang , Junghyun Kim , Kyuhwan Shim , Jun Ki Lee , Byoung-Tak Zhang

General robot skill adaptation requires expressive representations robust to varying task configurations. While recent learning-based skill adaptation methods refined via Reinforcement Learning (RL), have shown success, existing skill…

Collecting large amounts of real-world interaction data to train general robotic policies is often prohibitively expensive, thus motivating the use of simulation data. However, existing methods for data generation have generally focused on…

Machine Learning · Computer Science 2024-01-23 Lirui Wang , Yiyang Ling , Zhecheng Yuan , Mohit Shridhar , Chen Bao , Yuzhe Qin , Bailin Wang , Huazhe Xu , Xiaolong Wang

Visual actionable affordance has emerged as a transformative approach in robotics, focusing on perceiving interaction areas prior to manipulation. Traditional methods rely on pixel sampling to identify successful interaction samples or…

Robotics · Computer Science 2025-10-10 Taewhan Kim , Hojin Bae , Zeming Li , Xiaoqi Li , Iaroslav Ponomarenko , Ruihai Wu , Hao Dong

The acquisition of large-scale and diverse demonstration data are essential for improving robotic imitation learning generalization. However, generating such data for complex manipulations is challenging in real-world settings. We introduce…

Robotics · Computer Science 2025-03-18 Wensheng Wang , Ning Tan

Despite the potential of reinforcement learning (RL) for building general-purpose robotic systems, training RL agents to solve robotics tasks still remains challenging due to the difficulty of exploration in purely continuous action spaces.…

Machine Learning · Computer Science 2021-10-29 Murtaza Dalal , Deepak Pathak , Ruslan Salakhutdinov

Vision-Language-Action (VLA) models have emerged as a promising paradigm for robot learning, but their representations are still largely inherited from static image-text pretraining, leaving physical dynamics to be learned from…

Robotics · Computer Science 2026-03-24 Teli Ma , Jia Zheng , Zifan Wang , Chunli Jiang , Andy Cui , Junwei Liang , Shuo Yang