English
Related papers

Related papers: GigaBrain-0: A World Model-Powered Vision-Language…

200 papers

Recent Vision-Language-Action (VLA) models have made impressive progress toward general-purpose robotic manipulation by post-training large Vision-Language Models (VLMs) for action prediction. Yet most VLAs entangle perception and control…

Embodied AI is widely recognized as a cornerstone of artificial general intelligence (AGI) because it involves controlling embodied agents to perform tasks in the physical world. Building on the success of large language models (LLMs) and…

Robotics · Computer Science 2026-05-04 Yueen Ma , Zixing Song , Yuzheng Zhuang , Jianye Hao , Irwin King

Recent advancements in large multimodal models have led to the emergence of remarkable generalist capabilities in digital domains, yet their translation to physical agents such as robots remains a significant challenge. This report…

Robotics · Computer Science 2025-03-27 Gemini Robotics Team , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Travis Armstrong , Ashwin Balakrishna , Robert Baruch , Maria Bauza , Michiel Blokzijl , Steven Bohez , Konstantinos Bousmalis , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Oscar Chang , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , David D'Ambrosio , Sudeep Dasari , Todor Davchev , Coline Devin , Norman Di Palo , Tianli Ding , Adil Dostmohamed , Danny Driess , Yilun Du , Debidatta Dwibedi , Michael Elabd , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Leonard Hasenclever , Nicolas Heess , Brandon Hernaez , Alexander Herzog , R. Alex Hofer , Jan Humplik , Atil Iscen , Mithun George Jacob , Deepali Jain , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Chase Kew , Jerad Kirkland , Sean Kirmani , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Jacky Liang , Yixin Lin , Sharath Maddineni , Anirudha Majumdar , Assaf Hurwitz Michaely , Robert Moreno , Michael Neunert , Francesco Nori , Carolina Parada , Emilio Parisotto , Peter Pastor , Acorn Pooley , Kanishka Rao , Krista Reymann , Dorsa Sadigh , Stefano Saliceti , Pannag Sanketi , Pierre Sermanet , Dhruv Shah , Mohit Sharma , Kathryn Shea , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Jost Tobias Springenberg , Rachel Sterneck , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Vincent Vanhoucke , Jake Varley , Grace Vesom , Giulia Vezzani , Oriol Vinyals , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Yuxiang Yang , Rui Yao , Sergey Yaroshenko , Wenhao Yu , Wentao Yuan , Jingwei Zhang , Tingnan Zhang , Allan Zhou , Yuxiang Zhou

Vision-language-action models have advanced rapidly, but robot trajectories alone provide limited coverage for learning broad physical understanding. PhysBrain 1.0 studies a complementary route: converting large-scale human egocentric video…

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent…

The capability of performing long-horizon, language-guided robotic manipulation tasks critically relies on leveraging historical information and generating coherent action sequences. However, such capabilities are often overlooked by…

Robotics · Computer Science 2025-12-24 Xiaofan Wang , Xingyu Gao , Jianlong Fu , Zuolei Li , Dean Fortier , Galen Mullins , Andrey Kolobov , Baining Guo

The rapid progress of auto-regressive vision-language models (VLMs) has inspired growing interest in vision-language-action models (VLA) for robotic manipulation. Recently, masked diffusion models, a paradigm distinct from autoregressive…

Robotics · Computer Science 2025-09-11 Yuqing Wen , Hebei Li , Kefan Gu , Yucheng Zhao , Tiancai Wang , Xiaoyan Sun

Vision-Language-Action (VLA) models hold promise for generalist robotics but currently struggle with data scarcity, architectural inefficiencies, and the inability to generalize across different hardware platforms. We introduce RDT2, a…

Robotics · Computer Science 2026-02-04 Songming Liu , Bangguo Li , Kai Ma , Lingxuan Wu , Hengkai Tan , Xiao Ouyang , Hang Su , Jun Zhu

A robot in a human-centric environment needs to account for the human's intent and future motion in its task and motion planning to ensure safe and effective operation. This requires symbolic reasoning about probable future actions and the…

Robotics · Computer Science 2023-11-01 Moritz A. Graule , Volkan Isler

Building generalist embodied agents requires integrating perception, language understanding, and action, which are core capabilities addressed by Vision-Language-Action (VLA) approaches based on multimodal foundation models, including…

Robotics · Computer Science 2026-04-08 StarVLA Community

While Vision-Language-Action models (VLAs) are rapidly advancing towards generalist robot policies, it remains difficult to quantitatively understand their limits and failure modes. To address this, we introduce a comprehensive benchmark…

Robotics · Computer Science 2025-12-30 Borong Zhang , Jiahao Li , Jiachen Shen , Yishuai Cai , Yuhao Zhang , Yuanpei Chen , Juntao Dai , Jiaming Ji , Yaodong Yang

Training vision-based manipulation policies that are robust across diverse visual environments remains an important and unresolved challenge in robot learning. Current approaches often sidestep the problem by relying on invariant…

Robotics · Computer Science 2025-05-20 Sumeet Batra , Gaurav Sukhatme

While end-to-end Vision-Language-Action (VLA) models offer a promising paradigm for robotic manipulation, fine-tuning them on narrow control data often compromises the profound reasoning capabilities inherited from their base…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Tianshuo Yang , Guanyu Chen , Yutian Chen , Zhixuan Liang , Yitian Liu , Zanxin Chen , Chunpu Xu , Haotian Liang , Jiangmiao Pang , Yao Mu , Ping Luo

Vision-Language-Action (VLA) models hold great promise for general-purpose robotic intelligence, yet scaling up such models is severely bottlenecked by the high cost of acquiring annotated training data. Fortunately, vision-equipped robots…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Yuhao Zhou , Yunpeng Zhu , Yang Zhou , Jindi Lyu , Jian Lan , Zhangyuan Wang , Dan Si , Thomas Seidl , Qing Ye , Jiancheng Lyu

The integration of language instructions with robotic control, particularly through Vision Language Action (VLA) models, has shown significant potential. However, these systems are often hindered by high computational costs, the need for…

Robotics · Computer Science 2025-02-04 Marie Samson , Bastien Muraccioli , Fumio Kanehiro

Despite advances in Vision-Language-Action (VLA) models, robotic manipulation struggles with fine-grained tasks because current models lack mechanisms for active visual attention allocation. Human gaze naturally encodes intent, planning,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Anupam Pani , Yanchao Yang

Improving the generalization capabilities of general-purpose robotic manipulation agents in the real world has long been a significant challenge. Existing approaches often rely on collecting large-scale robotic data which is costly and…

Robotics · Computer Science 2025-02-10 Jiange Yang , Wenhui Tan , Chuhao Jin , Keling Yao , Bei Liu , Jianlong Fu , Ruihua Song , Gangshan Wu , Limin Wang

The emergence of Vision Language Action (VLA) models marks a paradigm shift from traditional policy-based control to generalized robotics, reframing Vision Language Models (VLMs) from passive sequence generators into active agents for…

Robotics · Computer Science 2025-11-11 Dapeng Zhang , Jing Sun , Chenghui Hu , Xiaoyan Wu , Zhenlong Yuan , Rui Zhou , Fei Shen , Qingguo Zhou

Trustworthy robot behavior requires not only high levels of task success but also that the robot can reliably quantify how likely it is to succeed. To this end, we present a first-of-its-kind study of confidence calibration in…

Robotics · Computer Science 2025-12-23 Thomas P Zollo , Richard Zemel

Vision-language-action (VLA) models show potential for general robotic tasks, but remain challenging in spatiotemporally coherent manipulation, which requires fine-grained representations. Typically, existing methods embed 3D positions into…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Hanyu Zhou , Chuanhao Ma , Gim Hee Lee
‹ Prev 1 8 9 10 Next ›