English
Related papers

Related papers: RLDX-1 Technical Report

200 papers

Vision-language-action models (VLAs) have garnered significant attention for their potential in advancing robotic manipulation. However, previous approaches predominantly rely on the general comprehension capabilities of vision-language…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Yuqi Wang , Xinghang Li , Wenxuan Wang , Junbo Zhang , Yingyan Li , Yuntao Chen , Xinlong Wang , Zhaoxiang Zhang

Vision-language-action (VLA) models finetuned from vision-language models (VLMs) hold the promise of leveraging rich pretrained representations to build generalist robots across diverse tasks and environments. However, direct fine-tuning on…

Robotics · Computer Science 2025-09-18 Shresth Grover , Akshay Gopalkrishnan , Bo Ai , Henrik I. Christensen , Hao Su , Xuanlin Li

The capability of performing long-horizon, language-guided robotic manipulation tasks critically relies on leveraging historical information and generating coherent action sequences. However, such capabilities are often overlooked by…

Robotics · Computer Science 2025-12-24 Xiaofan Wang , Xingyu Gao , Jianlong Fu , Zuolei Li , Dean Fortier , Galen Mullins , Andrey Kolobov , Baining Guo

Recently in robotics, Vision-Language-Action (VLA) models have emerged as a transformative approach, enabling robots to execute complex tasks by integrating visual and linguistic inputs within an end-to-end learning framework. Despite their…

Robotic manipulation in 3D requires effective computation of N degree-of-freedom joint-space trajectories that enable precise and robust control. To achieve this, robots must integrate semantic understanding with visual perception to…

Robotics · Computer Science 2026-03-31 Vineet Bhat , Yu-Hsiang Lan , Prashanth Krishnamurthy , Ramesh Karri , Farshad Khorrami

Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models (LLMs) and…

Robotics · Computer Science 2026-05-04 Yueen Ma , Zixing Song , Yuzheng Zhuang , Jianye Hao , Irwin King

We introduce OG-VLA, a novel architecture and learning framework that combines the generalization strengths of Vision Language Action models (VLAs) with the robustness of 3D-aware policies. We address the challenge of mapping natural…

Robotics · Computer Science 2025-11-19 Ishika Singh , Ankit Goyal , Stan Birchfield , Dieter Fox , Animesh Garg , Valts Blukis

Robotic real-world reinforcement learning (RL) with vision-language-action (VLA) models is bottlenecked by sparse, handcrafted rewards and inefficient exploration. We introduce VLAC, a general process reward model built upon InternVL and…

Dexterous grasping remains a fundamental yet challenging problem in robotics. A general-purpose robot must be capable of grasping diverse objects in arbitrary scenarios. However, existing research typically relies on restrictive…

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained…

Robotics · Computer Science 2025-12-09 Yichao Shen , Fangyun Wei , Zhiying Du , Yaobo Liang , Yan Lu , Jiaolong Yang , Nanning Zheng , Baining Guo

Although large vision-language-action (VLA) models pretrained on extensive robot datasets offer promising generalist policies for robotic learning, they still struggle with spatial-temporal dynamics in interactive robotics, making them less…

Vision-Language-Action (VLA) models have achieved remarkable success in robotic manipulation. However, their robustness to linguistic nuances remains a critical, under-explored safety concern, posing a significant safety risk to real-world…

Robotics · Computer Science 2026-04-08 Baoshun Tong , Haoran He , Ling Pan , Yang Liu , Liang Lin

Latent Action Models (LAMs) have emerged as an effective paradigm for handling heterogeneous datasets during Vision-Language-Action (VLA) model pretraining, offering a unified action space across embodiments. However, existing LAMs often…

Robotics · Computer Science 2026-05-14 Qiwei Li , Xicheng Gong , Xinghang Li , Peiyan Li , Quanyun Zhou , Hangjun Ye , Jiahuan Zhou , Yadong Mu

Vision-Language-Action (VLA) models have recently shown impressive generalization and language-guided manipulation capabilities. However, their performance degrades on tasks requiring precise spatial reasoning due to limited spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-10-16 Tianyuan Yuan , Yicheng Liu , Chenhao Lu , Zhuoguang Chen , Tao Jiang , Hang Zhao

The development of general robotic systems capable of manipulating in unstructured environments is a significant challenge. While Vision-Language Models(VLM) excel in high-level commonsense reasoning, they lack the fine-grained 3D spatial…

Robotics · Computer Science 2025-01-08 Mingjie Pan , Jiyao Zhang , Tianshu Wu , Yinghao Zhao , Wenlong Gao , Hao Dong

Vision-Language-Action (VLA) models have demonstrated remarkable capabilities in robotic manipulation,enabling robots to execute natural language commands through end-to-end learning from visual observations.However, deploying large-scale…

Robotics · Computer Science 2025-12-16 Abdullah Yahya Abdullah Omaisan , Ibrahim Sheikh Mohamed

Despite progress, Vision-Language-Action models (VLAs) are limited by a scarcity of large-scale, diverse robot data. While human manipulation videos offer a rich alternative, existing methods are forced to choose between small,…

Robotics · Computer Science 2026-02-26 Hao Luo , Ye Wang , Wanpeng Zhang , Haoqi Yuan , Yicheng Feng , Haiweng Xu , Sipeng Zheng , Zongqing Lu

The human ability to learn, generalize, and control complex manipulation tasks through multi-modality feedback suggests a unique capability, which we refer to as dexterity intelligence. Understanding and assessing this intelligence is a…

Robotics · Computer Science 2025-12-03 Fanlong Zeng , Wensheng Gan , Zezheng Huai , Lichao Sun , Hechang Chen , Yongheng Wang , Ning Liu , Philip S. Yu

Large, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with…

Robotics · Computer Science 2025-05-15 Embodiment Collaboration , Abby O'Neill , Abdul Rehman , Abhinav Gupta , Abhiram Maddukuri , Abhishek Gupta , Abhishek Padalkar , Abraham Lee , Acorn Pooley , Agrim Gupta , Ajay Mandlekar , Ajinkya Jain , Albert Tung , Alex Bewley , Alex Herzog , Alex Irpan , Alexander Khazatsky , Anant Rai , Anchit Gupta , Andrew Wang , Andrey Kolobov , Anikait Singh , Animesh Garg , Aniruddha Kembhavi , Annie Xie , Anthony Brohan , Antonin Raffin , Archit Sharma , Arefeh Yavary , Arhan Jain , Ashwin Balakrishna , Ayzaan Wahid , Ben Burgess-Limerick , Beomjoon Kim , Bernhard Schölkopf , Blake Wulfe , Brian Ichter , Cewu Lu , Charles Xu , Charlotte Le , Chelsea Finn , Chen Wang , Chenfeng Xu , Cheng Chi , Chenguang Huang , Christine Chan , Christopher Agia , Chuer Pan , Chuyuan Fu , Coline Devin , Danfei Xu , Daniel Morton , Danny Driess , Daphne Chen , Deepak Pathak , Dhruv Shah , Dieter Büchler , Dinesh Jayaraman , Dmitry Kalashnikov , Dorsa Sadigh , Edward Johns , Ethan Foster , Fangchen Liu , Federico Ceola , Fei Xia , Feiyu Zhao , Felipe Vieira Frujeri , Freek Stulp , Gaoyue Zhou , Gaurav S. Sukhatme , Gautam Salhotra , Ge Yan , Gilbert Feng , Giulio Schiavi , Glen Berseth , Gregory Kahn , Guangwen Yang , Guanzhi Wang , Hao Su , Hao-Shu Fang , Haochen Shi , Henghui Bao , Heni Ben Amor , Henrik I Christensen , Hiroki Furuta , Homanga Bharadhwaj , Homer Walke , Hongjie Fang , Huy Ha , Igor Mordatch , Ilija Radosavovic , Isabel Leal , Jacky Liang , Jad Abou-Chakra , Jaehyung Kim , Jaimyn Drake , Jan Peters , Jan Schneider , Jasmine Hsu , Jay Vakil , Jeannette Bohg , Jeffrey Bingham , Jeffrey Wu , Jensen Gao , Jiaheng Hu , Jiajun Wu , Jialin Wu , Jiankai Sun , Jianlan Luo , Jiayuan Gu , Jie Tan , Jihoon Oh , Jimmy Wu , Jingpei Lu , Jingyun Yang , Jitendra Malik , João Silvério , Joey Hejna , Jonathan Booher , Jonathan Tompson , Jonathan Yang , Jordi Salvador , Joseph J. Lim , Junhyek Han , Kaiyuan Wang , Kanishka Rao , Karl Pertsch , Karol Hausman , Keegan Go , Keerthana Gopalakrishnan , Ken Goldberg , Kendra Byrne , Kenneth Oslund , Kento Kawaharazuka , Kevin Black , Kevin Lin , Kevin Zhang , Kiana Ehsani , Kiran Lekkala , Kirsty Ellis , Krishan Rana , Krishnan Srinivasan , Kuan Fang , Kunal Pratap Singh , Kuo-Hao Zeng , Kyle Hatch , Kyle Hsu , Laurent Itti , Lawrence Yunliang Chen , Lerrel Pinto , Li Fei-Fei , Liam Tan , Linxi "Jim" Fan , Lionel Ott , Lisa Lee , Luca Weihs , Magnum Chen , Marion Lepert , Marius Memmel , Masayoshi Tomizuka , Masha Itkina , Mateo Guaman Castro , Max Spero , Maximilian Du , Michael Ahn , Michael C. Yip , Mingtong Zhang , Mingyu Ding , Minho Heo , Mohan Kumar Srirama , Mohit Sharma , Moo Jin Kim , Muhammad Zubair Irshad , Naoaki Kanazawa , Nicklas Hansen , Nicolas Heess , Nikhil J Joshi , Niko Suenderhauf , Ning Liu , Norman Di Palo , Nur Muhammad Mahi Shafiullah , Oier Mees , Oliver Kroemer , Osbert Bastani , Pannag R Sanketi , Patrick "Tree" Miller , Patrick Yin , Paul Wohlhart , Peng Xu , Peter David Fagan , Peter Mitrano , Pierre Sermanet , Pieter Abbeel , Priya Sundaresan , Qiuyu Chen , Quan Vuong , Rafael Rafailov , Ran Tian , Ria Doshi , Roberto Martín-Martín , Rohan Baijal , Rosario Scalise , Rose Hendrix , Roy Lin , Runjia Qian , Ruohan Zhang , Russell Mendonca , Rutav Shah , Ryan Hoque , Ryan Julian , Samuel Bustamante , Sean Kirmani , Sergey Levine , Shan Lin , Sherry Moore , Shikhar Bahl , Shivin Dass , Shubham Sonawani , Shubham Tulsiani , Shuran Song , Sichun Xu , Siddhant Haldar , Siddharth Karamcheti , Simeon Adebola , Simon Guist , Soroush Nasiriany , Stefan Schaal , Stefan Welker , Stephen Tian , Subramanian Ramamoorthy , Sudeep Dasari , Suneel Belkhale , Sungjae Park , Suraj Nair , Suvir Mirchandani , Takayuki Osa , Tanmay Gupta , Tatsuya Harada , Tatsuya Matsushima , Ted Xiao , Thomas Kollar , Tianhe Yu , Tianli Ding , Todor Davchev , Tony Z. Zhao , Travis Armstrong , Trevor Darrell , Trinity Chung , Vidhi Jain , Vikash Kumar , Vincent Vanhoucke , Vitor Guizilini , Wei Zhan , Wenxuan Zhou , Wolfram Burgard , Xi Chen , Xiangyu Chen , Xiaolong Wang , Xinghao Zhu , Xinyang Geng , Xiyuan Liu , Xu Liangwei , Xuanlin Li , Yansong Pang , Yao Lu , Yecheng Jason Ma , Yejin Kim , Yevgen Chebotar , Yifan Zhou , Yifeng Zhu , Yilin Wu , Ying Xu , Yixuan Wang , Yonatan Bisk , Yongqiang Dou , Yoonyoung Cho , Youngwoon Lee , Yuchen Cui , Yue Cao , Yueh-Hua Wu , Yujin Tang , Yuke Zhu , Yunchu Zhang , Yunfan Jiang , Yunshuang Li , Yunzhu Li , Yusuke Iwasawa , Yutaka Matsuo , Zehan Ma , Zhuo Xu , Zichen Jeff Cui , Zichen Zhang , Zipeng Fu , Zipeng Lin

Recent vision-language-action models (VLAs) build upon pretrained vision-language models and leverage diverse robot datasets to demonstrate strong task execution, language following ability, and semantic generalization. Despite these…

Robotics · Computer Science 2025-04-29 Moo Jin Kim , Chelsea Finn , Percy Liang
‹ Prev 1 4 5 6 7 8 10 Next ›