English
Related papers

Related papers: GR00T N1: An Open Foundation Model for Generalist …

200 papers

With the emergence of large language models (LLMs) and vision foundation models, how to combine the intelligence and capacity of these open-sourced or API-available models to achieve open-world visual perception remains an open question. In…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Chris Kelly , Luhui Hu , Bang Yang , Yu Tian , Deshun Yang , Cindy Yang , Zaoshan Huang , Zihao Li , Jiayin Hu , Yuexian Zou

This paper addresses the limitations of current humanoid robot control frameworks, which primarily rely on reactive mechanisms and lack autonomous interaction capabilities due to data scarcity. We propose Humanoid-VLA, a novel framework…

Generalization in robot manipulation is essential for deploying robots in open-world environments and advancing toward artificial general intelligence. While recent Vision-Language-Action (VLA) models leverage large pre-trained…

Robotics · Computer Science 2025-12-09 Yichao Shen , Fangyun Wei , Zhiying Du , Yaobo Liang , Yan Lu , Jiaolong Yang , Nanning Zheng , Baining Guo

Visual target navigation is a critical capability for autonomous robots operating in unknown environments, particularly in human-robot interaction scenarios. While classical and learning-based methods have shown promise, most existing…

Robotics · Computer Science 2025-05-07 Bangguo Yu , Qihao Yuan , Kailai Li , Hamidreza Kasaei , Ming Cao

Traditional robotic systems require specific training data for each task, environment, and robot form. While recent advancements in machine learning have enabled models to generalize across new tasks and environments, the challenge of…

Robotics · Computer Science 2024-09-06 Jonathan Salzer , Arnoud Visser

The rapid development of Large Language Models (LLMs) creates an exciting potential for flexible, general knowledge-driven Human-Robot Interaction (HRI) systems for assistive robots. Existing HRI systems demonstrate great progress in…

Robotics · Computer Science 2025-07-22 Jens V. Rüppel , Andrey Rudenko , Tim Schreiter , Martin Magnusson , Achim J. Lilienthal

Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as robots remains a significant challenge. This report…

Robotics · Computer Science 2025-03-27 Gemini Robotics Team , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Travis Armstrong , Ashwin Balakrishna , Robert Baruch , Maria Bauza , Michiel Blokzijl , Steven Bohez , Konstantinos Bousmalis , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Oscar Chang , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , David D'Ambrosio , Sudeep Dasari , Todor Davchev , Coline Devin , Norman Di Palo , Tianli Ding , Adil Dostmohamed , Danny Driess , Yilun Du , Debidatta Dwibedi , Michael Elabd , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Leonard Hasenclever , Nicolas Heess , Brandon Hernaez , Alexander Herzog , R. Alex Hofer , Jan Humplik , Atil Iscen , Mithun George Jacob , Deepali Jain , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Chase Kew , Jerad Kirkland , Sean Kirmani , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Jacky Liang , Yixin Lin , Sharath Maddineni , Anirudha Majumdar , Assaf Hurwitz Michaely , Robert Moreno , Michael Neunert , Francesco Nori , Carolina Parada , Emilio Parisotto , Peter Pastor , Acorn Pooley , Kanishka Rao , Krista Reymann , Dorsa Sadigh , Stefano Saliceti , Pannag Sanketi , Pierre Sermanet , Dhruv Shah , Mohit Sharma , Kathryn Shea , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Jost Tobias Springenberg , Rachel Sterneck , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Vincent Vanhoucke , Jake Varley , Grace Vesom , Giulia Vezzani , Oriol Vinyals , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Yuxiang Yang , Rui Yao , Sergey Yaroshenko , Wenhao Yu , Wentao Yuan , Jingwei Zhang , Tingnan Zhang , Allan Zhou , Yuxiang Zhou

General-purpose pre-trained models ("foundation models") have enabled practitioners to produce generalizable solutions for individual machine learning problems with datasets that are significantly smaller than those required for learning…

Robotics · Computer Science 2023-10-25 Dhruv Shah , Ajay Sridhar , Nitish Dashora , Kyle Stachowicz , Kevin Black , Noriaki Hirose , Sergey Levine

This paper presents a novel layered framework that integrates visual foundation models to improve robot manipulation tasks and motion planning. The framework consists of five layers: Perception, Cognition, Planning, Execution, and Learning.…

Robotics · Computer Science 2023-09-21 Chen Yang , Peng Zhou , Jiaming Qi

Large Language Models (LLMs) and strong vision models have enabled rapid research and development in the field of Vision-Language-Action models that enable robotic control. The main objective of these methods is to develop a generalist…

We present a new robotic foundation model, called ${\pi}_{0.7}$, that can enable strong out-of-the-box performance in a wide range of scenarios. ${\pi}_{0.7}$ can follow diverse language instructions in unseen environments, including…

The development of general robotic systems capable of manipulating in unstructured environments is a significant challenge. While Vision-Language Models(VLM) excel in high-level commonsense reasoning, they lack the fine-grained 3D spatial…

Robotics · Computer Science 2025-01-08 Mingjie Pan , Jiyao Zhang , Tianshu Wu , Yinghao Zhao , Wenlong Gao , Hao Dong

Vision-Language-Action (VLA) models demonstrate promising generalization in robotic manipulation, driven by advances in large-scale vision and language pre-training. This progress can be misleading. Despite the zero-shot perception and…

On March 18, 2024, NVIDIA unveiled Project GR00T, a general-purpose multimodal generative AI model designed specifically for training humanoid robots. Preceding this event, Tesla's unveiling of the Optimus Gen 2 humanoid robot on December…

Robotics · Computer Science 2024-04-01 Haiwei Dong , Yang Liu , Ted Chu , Abdulmotaleb El Saddik

The emerging field of Vision-Language-Action (VLA) for humanoid robots faces several fundamental challenges, including the high cost of data acquisition, the lack of a standardized benchmark, and the significant gap between simulation and…

Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully. Despite the success of large language models and vision language models based…

Artificial Intelligence · Computer Science 2025-10-23 Jinrui Liu , Bingyan Nie , Boyu Li , Yaran Chen , Yuze Wang , Shunsen He , Haoran Li

This article suggests a reasoning-guided vision-language-motion diffusion framework (RG-VLMD) for generating instruction-aware co-speech gestures for humanoid robots in educational scenarios. The system integrates multi-modal affective…

Robotics · Computer Science 2026-03-20 Fuze Sun , Lingyu Li , Lekan Dai , Xinyu Fan

Robotic manipulation faces a significant challenge in generalizing across unseen objects, environments and tasks specified by diverse language instructions. To improve generalization capabilities, recent research has incorporated large…

Robotics · Computer Science 2025-06-16 Shizhe Chen , Ricardo Garcia , Paul Pacaud , Cordelia Schmid

Multimodal models are expected to be a critical component to future advances in artificial intelligence. This field is starting to grow rapidly with a surge of new design elements motivated by the success of foundation models in natural…

Computation and Language · Computer Science 2024-06-11 Sai Munikoti , Ian Stewart , Sameera Horawalavithana , Henry Kvinge , Tegan Emerson , Sandra E Thompson , Karl Pazdernik

The advancement of large Vision-Language-Action (VLA) models has significantly improved robotic manipulation in terms of language-guided task execution and generalization to unseen scenarios. While existing VLAs adapted from pretrained…