中文
相关论文

相关论文: Unified Embodied VLM Reasoning with Robotic Action…

200 篇论文

Generalization in embodied AI is hindered by the "seeing-to-doing gap," which stems from data scarcity and embodiment heterogeneity. To address this, we pioneer "pointing" as a unified, embodiment-agnostic intermediate representation,…

机器人学 · 计算机科学 2026-04-07 Yifu Yuan , Haiqin Cui , Yaoting Huang , Yibin Chen , Fei Ni , Zibin Dong , Pengyi Li , Yan Zheng , Hongyao Tang , Jianye Hao

Embodied Question Answering (EQA) connects perception, reasoning, and interaction within embodied environments. However, existing datasets and benchmarks remain fragmented, each focusing on a limited subset of reasoning skills such as…

机器人学 · 计算机科学 2026-05-26 Xicheng Gong , Qiwei Li , Peiran Xu , Yadong Mu

Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and video understanding tasks. However, their effectiveness in embodied settings, which require online interaction and active scene…

计算机视觉与模式识别 · 计算机科学 2025-07-15 Mingxian Lin , Wei Huang , Yitang Li , Chengjie Jiang , Kui Wu , Fangwei Zhong , Shengju Qian , Xin Wang , Xiaojuan Qi

Recent Vision-Language-Action (VLA) models report impressive success rates on standard robotic benchmarks, fueling optimism about general-purpose physical intelligence. However, recent evidence suggests a systematic misalignment between…

机器人学 · 计算机科学 2026-04-21 Haiweng Xu , Sipeng Zheng , Hao Luo , Wanpeng Zhang , Ziheng Xi , Zongqing Lu

Improving the reasoning capabilities of embodied agents is crucial for robots to complete complex human instructions in long-view manipulation tasks successfully. Despite the success of large language models and vision language models based…

人工智能 · 计算机科学 2025-10-23 Jinrui Liu , Bingyan Nie , Boyu Li , Yaran Chen , Yuze Wang , Shunsen He , Haoran Li

A key limitation of learned robot control policies is their inability to generalize outside their training data. Recent works on vision-language-action models (VLAs) have shown that the use of large, internet pre-trained vision-language…

机器人学 · 计算机科学 2025-03-10 Michał Zawalski , William Chen , Karl Pertsch , Oier Mees , Chelsea Finn , Sergey Levine

While significant research has focused on developing embodied reasoning capabilities using Vision-Language Models (VLMs) or integrating advanced VLMs into Vision-Language-Action (VLA) models for end-to-end robot control, few studies…

Humans act with context and intention, with reasoning playing a central role. While internet-scale data has enabled broad reasoning capabilities in AI systems, grounding these abilities in physical action remains a major challenge. We…

机器人学 · 计算机科学 2025-12-11 Peijun Tang , Shangjin Xie , Binyan Sun , Baifu Huang , Kuncheng Luo , Haotian Yang , Weiqi Jin , Jianan Wang

Embodied robotic systems increasingly rely on large language model (LLM)-based agents to support high-level reasoning, planning, and decision-making during interactions with the environment. However, invoking LLM reasoning introduces…

Embodied agents operating in the physical world must make decisions that are not only effective but also safe, spatially coherent, and grounded in context. While recent advances in large multimodal models (LMMs) have shown promising…

General-purpose robots need a deep understanding of the physical world, advanced reasoning, and general and dexterous control. This report introduces the latest generation of the Gemini Robotics model family: Gemini Robotics 1.5, a…

机器人学 · 计算机科学 2025-12-02 Gemini Robotics Team , Abbas Abdolmaleki , Saminda Abeyruwan , Joshua Ainslie , Jean-Baptiste Alayrac , Montserrat Gonzalez Arenas , Ashwin Balakrishna , Nathan Batchelor , Alex Bewley , Jeff Bingham , Michael Bloesch , Konstantinos Bousmalis , Philemon Brakel , Anthony Brohan , Thomas Buschmann , Arunkumar Byravan , Serkan Cabi , Ken Caluwaerts , Federico Casarini , Christine Chan , Oscar Chang , London Chappellet-Volpini , Jose Enrique Chen , Xi Chen , Hao-Tien Lewis Chiang , Krzysztof Choromanski , Adrian Collister , David B. D'Ambrosio , Sudeep Dasari , Todor Davchev , Meet Kirankumar Dave , Coline Devin , Norman Di Palo , Tianli Ding , Carl Doersch , Adil Dostmohamed , Yilun Du , Debidatta Dwibedi , Sathish Thoppay Egambaram , Michael Elabd , Tom Erez , Xiaolin Fang , Claudio Fantacci , Cody Fong , Erik Frey , Chuyuan Fu , Ruiqi Gao , Marissa Giustina , Keerthana Gopalakrishnan , Laura Graesser , Oliver Groth , Agrim Gupta , Roland Hafner , Steven Hansen , Leonard Hasenclever , Sam Haves , Nicolas Heess , Brandon Hernaez , Alex Hofer , Jasmine Hsu , Lu Huang , Sandy H. Huang , Atil Iscen , Mithun George Jacob , Deepali Jain , Sally Jesmonth , Abhishek Jindal , Ryan Julian , Dmitry Kalashnikov , M. Emre Karagozler , Stefani Karp , Matija Kecman , J. Chase Kew , Donnie Kim , Frank Kim , Junkyung Kim , Thomas Kipf , Sean Kirmani , Ksenia Konyushkova , Li Yang Ku , Yuheng Kuang , Thomas Lampe , Antoine Laurens , Tuan Anh Le , Isabel Leal , Alex X. Lee , Tsang-Wei Edward Lee , Guy Lever , Jacky Liang , Li-Heng Lin , Fangchen Liu , Shangbang Long , Caden Lu , Sharath Maddineni , Anirudha Majumdar , Kevis-Kokitsi Maninis , Andrew Marmon , Sergio Martinez , Assaf Hurwitz Michaely , Niko Milonopoulos , Joss Moore , Robert Moreno , Michael Neunert , Francesco Nori , Joy Ortiz , Kenneth Oslund , Carolina Parada , Emilio Parisotto , Amaris Paryag , Acorn Pooley , Thomas Power , Alessio Quaglino , Haroon Qureshi , Rajkumar Vasudeva Raju , Helen Ran , Dushyant Rao , Kanishka Rao , Isaac Reid , David Rendleman , Krista Reymann , Miguel Rivas , Francesco Romano , Yulia Rubanova , Peter Pastor Sampedro , Pannag R Sanketi , Dhruv Shah , Mohit Sharma , Kathryn Shea , Mohit Shridhar , Charles Shu , Vikas Sindhwani , Sumeet Singh , Radu Soricut , Rachel Sterneck , Ian Storz , Razvan Surdulescu , Jie Tan , Jonathan Tompson , Saran Tunyasuvunakool , Jake Varley , Grace Vesom , Giulia Vezzani , Maria Bauza Villalonga , Oriol Vinyals , René Wagner , Ayzaan Wahid , Stefan Welker , Paul Wohlhart , Chengda Wu , Markus Wulfmeier , Fei Xia , Ted Xiao , Annie Xie , Jinyu Xie , Peng Xu , Sichun Xu , Ying Xu , Zhuo Xu , Jimmy Yan , Sherry Yang , Skye Yang , Yuxiang Yang , Hiu Hong Yu , Wenhao Yu , Wentao Yuan , Yuan Yuan , Jingwei Zhang , Tingnan Zhang , Zhiyuan Zhang , Allan Zhou , Guangyao Zhou , Yuxiang Zhou

Traditional reinforcement learning-based robotic control methods are often task-specific and fail to generalize across diverse environments or unseen objects and instructions. Visual Language Models (VLMs) demonstrate strong scene…

机器人学 · 计算机科学 2024-12-18 Qi Sun , Pengfei Hong , Tej Deep Pala , Vernon Toh , U-Xuan Tan , Deepanway Ghosal , Soujanya Poria

Large Vision-Language Models (LVLMs) have recently shown great promise in advancing robotics by combining embodied reasoning with robot control. A common approach involves training on embodied reasoning tasks related to robot control using…

机器人学 · 计算机科学 2026-01-19 Dongyoung Kim , Sumin Park , Huiwon Jang , Jinwoo Shin , Jaehyung Kim , Younggyo Seo

Coordinating multiple embodied agents in dynamic environments remains a core challenge in artificial intelligence, requiring both perception-driven reasoning and scalable cooperation strategies. While recent works have leveraged large…

人工智能 · 计算机科学 2026-01-23 Li Kang , Xiufeng Song , Heng Zhou , Yiran Qin , Jie Yang , Xiaohong Liu , Philip Torr , Lei Bai , Zhenfei Yin

Embodied Question Answering (EQA) is an essential yet challenging task for robot assistants. Large vision-language models (VLMs) have shown promise for EQA, but existing approaches either treat it as static video question answering without…

机器人学 · 计算机科学 2025-08-12 Kai Cheng , Zhengyuan Li , Xingpeng Sun , Byung-Cheol Min , Amrit Singh Bedi , Aniket Bera

Recent advances in deep thinking models have demonstrated remarkable reasoning capabilities on mathematical and coding tasks. However, their effectiveness in embodied domains which require continuous interaction with environments through…

Vision-language navigation requires agents to reason and act under constraints of embodiment. While vision-language models (VLMs) demonstrate strong generalization, current benchmarks provide limited understanding of how embodiment -- i.e.,…

机器人学 · 计算机科学 2025-12-23 Tin Stribor Sohn , Maximilian Dillitzer , Jason J. Corso , Eric Sax

Building robots that can perceive, reason, and act in dynamic, unstructured environments remains a core challenge. Recent embodied systems often adopt a dual-system paradigm, where System 2 handles high-level reasoning while System 1…

The realization of Artificial General Intelligence (AGI) necessitates Embodied AI agents capable of robust spatial perception, effective task planning, and adaptive execution in physical environments. However, current large language models…

Embodied intelligence is often studied through specialized models for individual tasks such as manipulation or navigation, resulting in fragmented capabilities and limited generalization across tasks, environments, and robot embodiments. In…

‹ 上一页 1 2 3 10 下一页 ›