English
Related papers

Related papers: GR00T N1: An Open Foundation Model for Generalist …

200 papers

Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precision. However, autonomous medical robotics has been limited by a fundamental data problem:…

Robotics · Computer Science 2026-04-30 Open-H-Embodiment Consortium , : , Nigel Nelson , Juo-Tung Chen , Jesse Haworth , Xinhao Chen , Lukas Zbinden , Dianye Huang , Alaa Eldin Abdelaal , Alberto Arezzo , Ayberk Acar , Farshid Alambeigi , Carlo Alberto Ammirati , Yunke Ao , Pablo David Aranda Rodriguez , Soofiyan Atar , Mattia Ballo , Noah Barnes , Federica Barontini , Filip Binkiewicz , Peter Black , Sebastian Bodenstedt , Leonardo Borgioli , Nikola Budjak , Benjamin Calmé , Fabio Carrillo , Nicola Cavalcanti , Changwei Chen , Haoxin Chen , Sihang Chen , Qihan Chen , Zhongyu Chen , Ziyang Chen , Shing Shin Cheng , Meiqing Cheng , Min Cheng , Zih-Yun Sarah Chiu , Xiangyu Chu , Camilo Correa-Gallego , Giulio Dagnino , Anton Deguet , Jacob Delgado , Jonathan C. DeLong , Kaizhong Deng , Alexander Dimitrakakis , Qingpeng Ding , Hao Ding , Giovanni Distefano , Daniel Donoho , Anqing Duan , Marco Esposito , Shane Farritor , Jad Fayad , Zahi Fayad , Mario Ferradosa , Filippo Filicori , Chelsea Finn , Philipp Fürnstahl , Jiawei Ge , Stamatia Giannarou , Xavier Giralt Ludevid , Frederic Giraud , Aditya Amit Godbole , Ken Goldberg , Antony Goldenberg , Diego Granero Marana , Xiaoqing Guo , Tamás Haidegger , Evan Hailey , Pascal Hansen , Ziyi Hao , Kush Hari , Kengo Hayashi , Jonathon Hawkins , Shelby Haworth , Ortrun Hellig , S. Duke Herrell , Zhouyang Hong , Andrew Howe , Junlei Hu , Zhaoyang Jacopo Hu , Ria Jain , Mohammad Rafiee Javazm , Howard Ji , Rui Ji , Jianmin Ji , Zhongliang Jiang , Dominic Jones , Jeffrey Jopling , Britton Jordan , Ran Ju , Michael Kam , Luoyao Kang , Fausto Kang , Siddhartha Kapuria , Peter Kazanzides , Sonika Kiehler , Ethan Kilmer , Ji Woong Kim , Przemysław Korzeniowski , Chandra Kuchi , Nithesh Kumar , Alan Kuntz , Federico Lavagno , Yu Chung Lee , Hao-Chih Lee , Hang Li , Zhen Li , Xiao Liang , Xinxin Lin , Jinsong Lin , Chang Liu , Fei Liu , Pei Liu , Yun-hui Liu , Wanli Liuchen , Eszter Lukács , Sareena Mann , Miles Mannas , Brett Marinelli , Sabina Martyniak , Francesco Marzola , Lorenzo Mazza , Xueyan Mei , Maria Clara Morais , Luigi Muratore , Chetan Reddy Narayanaswamy , Michał Naskręt , David Navarro-Alarcon , Cyrus Neary , Chi Kit Ng , Christopher Nguan , David Noonan , Ki Hwan Oh , Tom Christian Olesch , Allison M. Okamura , Justin Opfermann , Matteo Pescio , Doan Xuan Viet Pham , Tito Porras , Hongliang Ren , Ariel Rodriguez Jimenez , Ferdinando Rodriguez y Baena , Septimiu E. Salcudean , Asmitha Sathya , Preethi Satish , Lalithkumar Seenivasan , Jiaqi Shao , Yiqing Shen , Yu Sheng , Lucy XiaoYang Shi , Zoe Soulé , Stefanie Speidel , Mingwu Su , Jianhao Su , Idris Sunmola , Kristóf Takács , Yunxi Tang , Patrick Thornycroft , Yu Tian , Jordan Thompson , Mehmet K. Turkcan , Mathias Unberath , Pietro Valdastri , Carlos Vives , Quan Vuong , Martin Wagner , Farong Wang , Wei Wang , Lidian Wang , Chung-Pang Wang , Guankun Wang , Junyi Wang , Erqi Wang , Ziyi Wang , Tanner Watts , Wolfgang Wein , Yimeng Wu , Zijian Wu , Hongjun Wu , Luohong Wu , Jie Ying Wu , Junlin Wu , Victoria Wu , Kaixuan Wu , Mateusz Wójcikowski , Yunye Xiao , Nan Xiao , Wenxuan Xie , Hao Yang , Tianqi Yang , Yinuo Yang , Menglong Ye , Ryan S. Yeung , Nural Yilmaz , Chim Ho Yin , Michael Yip , Rayan Younis , Chenhao Yu , Sayem Nazmuz Zaman , Milos Zefran , Han Zhang , Yuelin Zhang , Yidong Zhang , Yanyong Zhang , Xuyang Zhang , Yameng Zhang , Joyce Zhang , Ning Zhong , Peng Zhou , Haoying Zhou , Xiuli Zuo , Nassir Navab , Mahdi Azizian , Sean D. Huver , Axel Krieger

This paper presents RynnVLA-001, a vision-language-action(VLA) model built upon large-scale video generative pretraining from human demonstrations. We propose a novel two-stage pretraining methodology. The first stage, Ego-Centric Video…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Yuming Jiang , Siteng Huang , Shengke Xue , Yaxi Zhao , Jun Cen , Sicong Leng , Kehan Li , Jiayan Guo , Kexiang Wang , Mingxiu Chen , Fan Wang , Deli Zhao , Xin Li

Vision-Language-Action (VLA) models have recently become highly prominent in the field of robotics. Leveraging vision-language foundation models trained on large-scale internet data, the VLA model can generate robotic actions directly from…

Robotics · Computer Science 2025-05-19 Wei Zhao , Gongsheng Li , Zhefei Gong , Pengxiang Ding , Han Zhao , Donglin Wang

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the Gemini Robotics model family: Gemini Robotics 1.5, a…

Robotics · Computer Science 2025-12-02 Gemini Robotics Team , Abbas Abdolmaleki , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Ashwin Balakrishna , Nathan Batchelor , Alex Bewley , Jeff Bingham , Michael Bloesch , Konstantinos Bousmalis , Philemon Brakel , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Christine Chan , Oscar Chang , London Chappellet-Volpini , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , Adrian Collister , David B. D'Ambrosio , Sudeep Dasari , Todor Davchev , Meet Kirankumar Dave , Coline Devin , Norman Di Palo , Tianli Ding , Carl Doersch , Adil Dostmohamed , Yilun Du , Debidatta Dwibedi , Sathish Thoppay Egambaram , Michael Elabd , Tom Erez , Xiaolin Fang , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Ruiqi Gao , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Oliver Groth , Agrim Gupta , Roland Hafner , Steven Hansen , Leonard Hasenclever , Sam Haves , Nicolas Heess , Brandon Hernaez , Alex Hofer , Jasmine Hsu , Lu Huang , Sandy H. Huang , Atil Iscen , Mithun George Jacob , Deepali Jain , Sally Jesmonth , Abhishek Jindal , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Matija Kecman , J. Chase Kew , Donnie Kim , Frank Kim , Junkyung Kim , Thomas Kipf , Sean Kirmani , Ksenia Konyushkova , Li Yang Ku , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Tuan Anh Le , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Guy Lever , Jacky Liang , Li-Heng Lin , Fangchen Liu , Shangbang Long , Caden Lu , Sharath Maddineni , Anirudha Majumdar , Kevis-Kokitsi Maninis , Andrew Marmon , Sergio Martinez , Assaf Hurwitz Michaely , Niko Milonopoulos , Joss Moore , Robert Moreno , Michael Neunert , Francesco Nori , Joy Ortiz , Kenneth Oslund , Carolina Parada , Emilio Parisotto , Amaris Paryag , Acorn Pooley , Thomas Power , Alessio Quaglino , Haroon Qureshi , Rajkumar Vasudeva Raju , Helen Ran , Dushyant Rao , Kanishka Rao , Isaac Reid , David Rendleman , Krista Reymann , Miguel Rivas , Francesco Romano , Yulia Rubanova , Peter Pastor Sampedro , Pannag R Sanketi , Dhruv Shah , Mohit Sharma , Kathryn Shea , Mohit Shridhar , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Rachel Sterneck , Ian Storz , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Saran Tunyasuvunakool , Jake Varley , Grace Vesom , Giulia Vezzani , Maria Bauza Villalonga , Oriol Vinyals , René Wagner , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Chengda Wu , Markus Wulfmeier , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Jimmy Yan , Sherry Yang , Skye Yang , Yuxiang Yang , Hiu Hong Yu , Wenhao Yu , Wentao Yuan , Yuan Yuan , Jingwei Zhang , Tingnan Zhang , Zhiyuan Zhang , Allan Zhou , Guangyao Zhou , Yuxiang Zhou

We pursue the goal of developing robots that can interact zero-shot with generic unseen objects via a diverse repertoire of manipulation skills and show how passive human videos can serve as a rich source of data for learning such…

Robotics · Computer Science 2023-12-04 Homanga Bharadhwaj , Abhinav Gupta , Vikash Kumar , Shubham Tulsiani

World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA)…

We present GR-RL, a robotic learning framework that turns a generalist vision-language-action (VLA) policy into a highly capable specialist for long-horizon dexterous manipulation. Assuming the optimality of human demonstrations is core to…

Large policies pretrained on a combination of Internet-scale vision-language data and diverse robot demonstrations have the potential to change how we teach robots new skills: rather than training new behaviors from scratch, we can…

One central goal of robotics is to enable robots to interact with the physical world. Traditional manipulation studies primarily focus on single robots and relatively small objects. However, factory and domestic environments often require…

Robotics · Computer Science 2026-05-26 Kun Song , Gaoming Chen , Shentao Ma , Ninglong Jin , Guangbao Zhao , Mingyu Ding , Zhenhua Xiong , Jia Pan

A key challenge in robotic manipulation in open domains is how to acquire diverse and generalizable skills for robots. Recent research in one-shot imitation learning has shown promise in transferring trained policies to new tasks based on…

Robotics · Computer Science 2023-09-27 Hao-Shu Fang , Hongjie Fang , Zhenyu Tang , Jirong Liu , Chenxi Wang , Junbo Wang , Haoyi Zhu , Cewu Lu

Towards the role of humanoid robots as squad mates in urban operations and other domains, we identified doors as a major area lacking capability development. In this paper, we focus on the ability of humanoid robots to navigate and deal…

To utilize Foundation Vision Language Models (VLMs) for robotic tasks and motion planning, the community has proposed different methods for injecting action components into VLMs and building the Vision-Language-Action models (VLAs). In this…

This paper presents GenH2R, a framework for learning generalizable vision-based human-to-robot (H2R) handover skills. The goal is to equip robots with the ability to reliably receive objects with unseen geometry handed over by humans in…

Robotics · Computer Science 2024-06-17 Zifan Wang , Junyu Chen , Ziqing Chen , Pengwei Xie , Rui Chen , Li Yi

We address the challenge of acquiring real-world manipulation skills with a scalable framework. We hold the belief that identifying an appropriate prediction target capable of leveraging large-scale datasets is crucial for achieving…

Robotics · Computer Science 2024-09-24 Chengbo Yuan , Chuan Wen , Tong Zhang , Yang Gao

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal,…

Artificial Intelligence · Computer Science 2025-05-21 Joel Currie , Gioele Migno , Enrico Piacenti , Maria Elena Giannaccini , Patric Bach , Davide De Tommaso , Agnieszka Wykowska

Visual Navigation Models (VNMs) promise generalizable, robot navigation by learning from large-scale visual demonstrations. Despite growing real-world deployment, existing evaluations rely almost exclusively on success rate, whether the…

Robotics · Computer Science 2026-03-30 Maeva Guerrier , Karthik Soma , Jana Pavlasek , Giovanni Beltrame

Recent robot learning methods commonly rely on imitation learning from massive robotic dataset collected with teleoperation. When facing a new task, such methods generally require collecting a set of new teleoperation data and finetuning…

Robotics · Computer Science 2025-05-28 Xiang Zhu , Yichen Liu , Hezhong Li , Jianyu Chen

While Vision-Language-Action (VLA) models show strong generalizability in various tasks, real-world deployment of robotic policy still requires large-scale, high-quality human expert demonstrations. However, data collection via human…

It is a long-standing problem in robotics to develop agents capable of executing diverse manipulation tasks from visual observations in unstructured real-world environments. To achieve this goal, the robot needs to have a comprehensive…

We propose LEO-RobotAgent, a general-purpose language-driven intelligent agent framework for robots. Under this framework, LLMs can operate different types of robots to complete unpredictable complex tasks across various scenarios. This…

Robotics · Computer Science 2026-04-16 Lihuang Chen , Xiangyu Luo , Jun Meng