English
Related papers

Related papers: Being-H0.5: Scaling Human-Centric Robot Learning f…

200 papers

The human ability to seamlessly perform multimodal reasoning and physical interaction in the open world is a core goal for general purpose embodied intelligent systems. Recent vision-language-action (VLA) models, which are co-trained on…

Existing Visual-Language-Action (VLA) models have shown promising performance in zero-shot scenarios, demonstrating impressive task execution and reasoning capabilities. However, a significant challenge arises from the limitations of visual…

Robotics · Computer Science 2025-04-29 Chia-Yu Hung , Qi Sun , Pengfei Hong , Amir Zadeh , Chuan Li , U-Xuan Tan , Navonil Majumder , Soujanya Poria

Long-term Human-Robot Collaboration (HRC) is crucial for enabling flexible manufacturing systems and integrating companion robots into daily human environments over extended periods. This paper identifies several key challenges for such…

Robotics · Computer Science 2025-02-05 Peiqi Yu , Abulikemu Abuduweili , Ruixuan Liu , Changliu Liu

Cross-embodiment learning from human demonstrations is hindered by the visual gap between human and robot embodiments. While self-supervised learning (SSL) backbones encode rich inter-class semantics of general objects, we show they fail to…

Vision-Language-Action (VLA) models have shown strong potential for general-purpose robot manipulation by unifying perception and action. However, existing VLA systems primarily rely on textual instructions and struggle to resolve spatial…

Robotics · Computer Science 2026-05-22 Wenxuan Guo , Ziyuan Li , Meng Zhang , Yichen Liu , Yimeng Dong , Chuxi Xu , Yunfei Wei , Ze Chen , Erjin Zhou , Jianjiang Feng

In recent years, instruction-tuned Large Multimodal Models (LMMs) have been successful at several tasks, including image captioning and visual question answering; yet leveraging these models remains an open question for robotics. Prior LMMs…

The remarkable progress of Multimodal Large Language Models (MLLMs) has attracted increasing attention to extend them to physical entities like legged robot. This typically requires MLLMs to not only grasp multimodal understanding…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Gen Luo , Ganlin Yang , Ziyang Gong , Guanzhou Chen , Haonan Duan , Erfei Cui , Ronglei Tong , Zhi Hou , Tianyi Zhang , Zhe Chen , Shenglong Ye , Lewei Lu , Jingbo Wang , Wenhai Wang , Jifeng Dai , Yu Qiao , Rongrong Ji , Xizhou Zhu

Vision-Language-Action (VLA) models are emerging as a promising paradigm for end-to-end autonomous driving, valued for their potential to leverage world knowledge and reason about complex driving scenes. However, existing methods suffer…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Xinyang Wang , Qian Liu , Wenjie Ding , Zhao Yang , Wei Li , Chang Liu , Bailin Li , Kun Zhan , Xianpeng Lang , Wei Chen

Training manipulation policies for humanoid robots with diverse data enhances their robustness and generalization across tasks and platforms. However, learning solely from robot demonstrations is labor-intensive, requiring expensive…

This paper introduces a new hybrid framework that combines Reinforcement Learning (RL) and Large Language Models (LLMs) to improve robotic manipulation tasks. By utilizing RL for accurate low-level control and LLMs for high level task…

Robotics · Computer Science 2026-04-01 Md Saad , Sajjad Hussain , Mohd Suhaib

Training on diverse, internet-scale data is a key factor in the success of recent large foundation models. Yet, using the same recipe for building embodied agents has faced noticeable difficulties. Despite the availability of many…

Vision-Language-Action (VLA) models have emerged as a powerful framework that unifies perception, language, and control, enabling robots to perform diverse tasks through multimodal understanding. However, current VLA models typically…

Vision-Language-Action (VLA) models typically bridge the gap between perceptual and action spaces by pre-training a large-scale Vision-Language Model (VLM) on robotic data. While this approach greatly enhances performance, it also incurs…

Achieving human-like dexterous manipulation remains a major challenge for general-purpose robots. While Vision-Language-Action (VLA) models show potential in learning skills from demonstrations, their scalability is limited by scarce…

Robotics · Computer Science 2025-12-16 Yu Cui , Yujian Zhang , Lina Tao , Yang Li , Xinyu Yi , Zhibin Li

Cross-embodiment manipulation is crucial for enhancing the scalability of robot manipulation and reducing the high cost of data collection. However, the significant differences between embodiments, such as variations in action spaces and…

Robotics · Computer Science 2026-03-17 Juncheng Mu , Sizhe Yang , Hojin Bae , Feiyu Jia , Qingwei Ben , Boyi Li , Huazhe Xu , Jiangmiao Pang

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the Gemini Robotics model family: Gemini Robotics 1.5, a…

Robotics · Computer Science 2025-12-02 Gemini Robotics Team , Abbas Abdolmaleki , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Ashwin Balakrishna , Nathan Batchelor , Alex Bewley , Jeff Bingham , Michael Bloesch , Konstantinos Bousmalis , Philemon Brakel , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Christine Chan , Oscar Chang , London Chappellet-Volpini , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , Adrian Collister , David B. D'Ambrosio , Sudeep Dasari , Todor Davchev , Meet Kirankumar Dave , Coline Devin , Norman Di Palo , Tianli Ding , Carl Doersch , Adil Dostmohamed , Yilun Du , Debidatta Dwibedi , Sathish Thoppay Egambaram , Michael Elabd , Tom Erez , Xiaolin Fang , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Ruiqi Gao , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Oliver Groth , Agrim Gupta , Roland Hafner , Steven Hansen , Leonard Hasenclever , Sam Haves , Nicolas Heess , Brandon Hernaez , Alex Hofer , Jasmine Hsu , Lu Huang , Sandy H. Huang , Atil Iscen , Mithun George Jacob , Deepali Jain , Sally Jesmonth , Abhishek Jindal , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Matija Kecman , J. Chase Kew , Donnie Kim , Frank Kim , Junkyung Kim , Thomas Kipf , Sean Kirmani , Ksenia Konyushkova , Li Yang Ku , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Tuan Anh Le , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Guy Lever , Jacky Liang , Li-Heng Lin , Fangchen Liu , Shangbang Long , Caden Lu , Sharath Maddineni , Anirudha Majumdar , Kevis-Kokitsi Maninis , Andrew Marmon , Sergio Martinez , Assaf Hurwitz Michaely , Niko Milonopoulos , Joss Moore , Robert Moreno , Michael Neunert , Francesco Nori , Joy Ortiz , Kenneth Oslund , Carolina Parada , Emilio Parisotto , Amaris Paryag , Acorn Pooley , Thomas Power , Alessio Quaglino , Haroon Qureshi , Rajkumar Vasudeva Raju , Helen Ran , Dushyant Rao , Kanishka Rao , Isaac Reid , David Rendleman , Krista Reymann , Miguel Rivas , Francesco Romano , Yulia Rubanova , Peter Pastor Sampedro , Pannag R Sanketi , Dhruv Shah , Mohit Sharma , Kathryn Shea , Mohit Shridhar , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Rachel Sterneck , Ian Storz , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Saran Tunyasuvunakool , Jake Varley , Grace Vesom , Giulia Vezzani , Maria Bauza Villalonga , Oriol Vinyals , René Wagner , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Chengda Wu , Markus Wulfmeier , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Jimmy Yan , Sherry Yang , Skye Yang , Yuxiang Yang , Hiu Hong Yu , Wenhao Yu , Wentao Yuan , Yuan Yuan , Jingwei Zhang , Tingnan Zhang , Zhiyuan Zhang , Allan Zhou , Guangyao Zhou , Yuxiang Zhou

In this study, we are interested in imbuing robots with the capability of physically-grounded task planning. Recent advancements have shown that large language models (LLMs) possess extensive knowledge useful in robotic tasks, especially in…

Robotics · Computer Science 2023-12-27 Yingdong Hu , Fanqi Lin , Tong Zhang , Li Yi , Yang Gao

Vision-Language-Action (VLA) models often suffer from performance degradation under distribution shifts, as they struggle to learn generalized behavior representations across varying environments. While existing approaches attempt to…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Bing Hu , Zaijing Li , Rui Shao , Junda Chen , April Hua Liu , Wei-Shi Zheng , Liqiang Nie

Robotic manipulation in 3D requires effective computation of N degree-of-freedom joint-space trajectories that enable precise and robust control. To achieve this, robots must integrate semantic understanding with visual perception to…

Robotics · Computer Science 2026-03-31 Vineet Bhat , Yu-Hsiang Lan , Prashanth Krishnamurthy , Ramesh Karri , Farshad Khorrami

Learning from real-world robot demonstrations holds promise for interacting with complex real-world environments. However, the complexity and variability of interaction dynamics often cause purely positional controllers to struggle with…

Robotics · Computer Science 2025-11-19 Lai Wei , Xuanbin Peng , Ri-Zhao Qiu , Tianshu Huang , Xuxin Cheng , Xiaolong Wang