English
Related papers

Related papers: EO-1: An Open Unified Embodied Foundation Model fo…

200 papers

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the Gemini Robotics model family: Gemini Robotics 1.5, a…

Robotics · Computer Science 2025-12-02 Gemini Robotics Team , Abbas Abdolmaleki , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Ashwin Balakrishna , Nathan Batchelor , Alex Bewley , Jeff Bingham , Michael Bloesch , Konstantinos Bousmalis , Philemon Brakel , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Christine Chan , Oscar Chang , London Chappellet-Volpini , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , Adrian Collister , David B. D'Ambrosio , Sudeep Dasari , Todor Davchev , Meet Kirankumar Dave , Coline Devin , Norman Di Palo , Tianli Ding , Carl Doersch , Adil Dostmohamed , Yilun Du , Debidatta Dwibedi , Sathish Thoppay Egambaram , Michael Elabd , Tom Erez , Xiaolin Fang , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Ruiqi Gao , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Oliver Groth , Agrim Gupta , Roland Hafner , Steven Hansen , Leonard Hasenclever , Sam Haves , Nicolas Heess , Brandon Hernaez , Alex Hofer , Jasmine Hsu , Lu Huang , Sandy H. Huang , Atil Iscen , Mithun George Jacob , Deepali Jain , Sally Jesmonth , Abhishek Jindal , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Matija Kecman , J. Chase Kew , Donnie Kim , Frank Kim , Junkyung Kim , Thomas Kipf , Sean Kirmani , Ksenia Konyushkova , Li Yang Ku , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Tuan Anh Le , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Guy Lever , Jacky Liang , Li-Heng Lin , Fangchen Liu , Shangbang Long , Caden Lu , Sharath Maddineni , Anirudha Majumdar , Kevis-Kokitsi Maninis , Andrew Marmon , Sergio Martinez , Assaf Hurwitz Michaely , Niko Milonopoulos , Joss Moore , Robert Moreno , Michael Neunert , Francesco Nori , Joy Ortiz , Kenneth Oslund , Carolina Parada , Emilio Parisotto , Amaris Paryag , Acorn Pooley , Thomas Power , Alessio Quaglino , Haroon Qureshi , Rajkumar Vasudeva Raju , Helen Ran , Dushyant Rao , Kanishka Rao , Isaac Reid , David Rendleman , Krista Reymann , Miguel Rivas , Francesco Romano , Yulia Rubanova , Peter Pastor Sampedro , Pannag R Sanketi , Dhruv Shah , Mohit Sharma , Kathryn Shea , Mohit Shridhar , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Rachel Sterneck , Ian Storz , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Saran Tunyasuvunakool , Jake Varley , Grace Vesom , Giulia Vezzani , Maria Bauza Villalonga , Oriol Vinyals , René Wagner , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Chengda Wu , Markus Wulfmeier , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Jimmy Yan , Sherry Yang , Skye Yang , Yuxiang Yang , Hiu Hong Yu , Wenhao Yu , Wentao Yuan , Yuan Yuan , Jingwei Zhang , Tingnan Zhang , Zhiyuan Zhang , Allan Zhou , Guangyao Zhou , Yuxiang Zhou

To operate effectively in the real world, robots should integrate multimodal reasoning with precise action generation. However, existing vision-language-action (VLA) models often sacrifice one for the other, narrow their abilities to…

Robotics · Computer Science 2026-03-04 Shuai Yang , Hao Li , Bin Wang , Yilun Chen , Yang Tian , Tai Wang , Hanqing Wang , Feng Zhao , Yiyi Liao , Jiangmiao Pang

The remarkable success of multimodal large language models (MLLMs) has driven advances in multimodal embeddings, yet existing models remain inherently discriminative, limiting their ability to benefit from reasoning-driven generation…

Machine Learning · Computer Science 2026-03-03 Zhibin Lan , Liqiang Niu , Fandong Meng , Jie Zhou , Jinsong Su

Enabling robots to understand language instructions and react accordingly to visual perception has been a long-standing goal in the robotics research community. Achieving this goal requires cutting-edge advances in natural language…

Robotics · Computer Science 2023-09-01 Tianyu Wang , Yifan Li , Haitao Lin , Xiangyang Xue , Yanwei Fu

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a unified framework.…

Robotics · Computer Science 2026-03-25 Shaid Hasan , Breenice Lee , Sujan Sarker , Tariq Iqbal

Vision-Language-Action models have emerged as a promising paradigm for robotic manipulation by unifying perception, language grounding, and action generation. However, they often struggle in scenarios requiring precise spatial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Tao Lin , Yuxin Du , Jiting Liu , Nuobei Zhu , Yunhe Li , Yuqian Fu , Yinxinyu Chen , Hongyi Cai , Zewei Ye , Bing Cheng , Kai Ye , Yiran Mao , Yilei Zhong , MingKang Dong , Junchi Yan , Gen Li , Bo Zhao

Multimodal Large Language Models have demonstrated powerful cross-modal understanding and reasoning capabilities in general domains. However, in the electromagnetic (EM) domain, they still face challenges such as data scarcity and…

Embodied robots which can interact with their environment and neighbours are increasingly being used as a test case to develop Artificial Intelligence. This creates a need for multimodal robot controllers that can operate across different…

Robotics · Computer Science 2025-02-05 William Hunt , Sarvapali D. Ramchurn , Mohammad D. Soorati

End-to-end architectures trained via imitation learning have advanced autonomous driving by scaling model size and data, yet performance remains brittle in safety-critical long-tail scenarios where supervision is sparse and causal…

Vision-language navigation requires agents to reason and act under constraints of embodiment. While vision-language models (VLMs) demonstrate strong generalization, current benchmarks provide limited understanding of how embodiment -- i.e.,…

Robotics · Computer Science 2025-12-23 Tin Stribor Sohn , Maximilian Dillitzer , Jason J. Corso , Eric Sax

Recent advances in deep thinking models have demonstrated remarkable reasoning capabilities on mathematical and coding tasks. However, their effectiveness in embodied domains which require continuous interaction with environments through…

Computation and Language · Computer Science 2025-05-15 Wenqi Zhang , Mengna Wang , Gangao Liu , Xu Huixin , Yiwei Jiang , Yongliang Shen , Guiyang Hou , Zhe Zheng , Hang Zhang , Xin Li , Weiming Lu , Peng Li , Yueting Zhuang

Vision-Language-Action (VLA) models benefit from chain-of-thought (CoT) reasoning, but existing approaches incur high inference overhead and rely on discrete reasoning representations that mismatch continuous perception and control. We…

We propose LEO-RobotAgent, a general-purpose language-driven intelligent agent framework for robots. Under this framework, LLMs can operate different types of robots to complete unpredictable complex tasks across various scenarios. This…

Robotics · Computer Science 2026-04-16 Lihuang Chen , Xiangyu Luo , Jun Meng

General-purpose robots capable of performing diverse tasks require synergistic reasoning and acting capabilities. However, recent dual-system approaches, which separate high-level reasoning from low-level acting, often suffer from…

Robotics · Computer Science 2026-03-03 Fanqi Lin , Ruiqian Nai , Yingdong Hu , Jiacheng You , Junming Zhao , Yang Gao

Human demonstrations offer rich environmental diversity and scale naturally, making them an appealing alternative to robot teleoperation. While this paradigm has advanced robot-arm manipulation, its potential for the more challenging,…

Robotics · Computer Science 2026-02-11 Modi Shi , Shijia Peng , Jin Chen , Haoran Jiang , Yinghui Li , Di Huang , Ping Luo , Hongyang Li , Li Chen

Humanoid robots are capable of performing various actions such as greeting, dancing and even backflipping. However, these motions are often hard-coded or specifically trained, which limits their versatility. In this work, we present…

We present a scalable framework for cross-embodiment humanoid robot control by learning a shared latent representation that unifies motion across humans and diverse humanoid platforms, including single-arm, dual-arm, and legged humanoid…

Robotics · Computer Science 2026-01-23 Yashuai Yan , Dongheui Lee

While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (VLA) models for end-to-end robot control, few studies…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Ganlin Yang , Tianyi Zhang , Haoran Hao , Weiyun Wang , Yibin Liu , Dehui Wang , Guanzhou Chen , Zijian Cai , Junting Chen , Weijie Su , Wengang Zhou , Yu Qiao , Jifeng Dai , Jiangmiao Pang , Gen Luo , Wenhai Wang , Yao Mu , Zhi Hou

Robotic foundation models trained on large-scale manipulation datasets have shown promise in learning generalist policies, but they often overfit to specific viewpoints, robot arms, and especially parallel-jaw grippers due to dataset…

Robotics · Computer Science 2026-01-15 Tong Wu , Shoujie Li , Junhao Gong , Changqing Guo , Xingting Li , Shilong Mu , Wenbo Ding

Multimodal Large Language Models (MLLMs) are making significant progress in multimodal reasoning. Early approaches focus on pure text-based reasoning. More recent studies have incorporated multimodal information into the reasoning steps;…

Artificial Intelligence · Computer Science 2026-04-21 Dongjie Cheng , Yongqi Li , Zhixin Ma , Hongru Cai , Yupeng Hu , Wenjie Wang , Liqiang Nie , Wenjie Li
‹ Prev 1 3 4 5 6 7 10 Next ›