English
Related papers

Related papers: EO-1: An Open Unified Embodied Foundation Model fo…

200 papers

Foundation models have revolutionized robotics by providing rich semantic representations without task-specific training. While many approaches integrate pretrained vision-language models (VLMs) with specialized navigation architectures,…

Autonomous medical robots hold promise to improve patient outcomes, reduce provider workload, democratize access to care, and enable superhuman precision. However, autonomous medical robotics has been limited by a fundamental data problem:…

Robotics · Computer Science 2026-04-30 Open-H-Embodiment Consortium , : , Nigel Nelson , Juo-Tung Chen , Jesse Haworth , Xinhao Chen , Lukas Zbinden , Dianye Huang , Alaa Eldin Abdelaal , Alberto Arezzo , Ayberk Acar , Farshid Alambeigi , Carlo Alberto Ammirati , Yunke Ao , Pablo David Aranda Rodriguez , Soofiyan Atar , Mattia Ballo , Noah Barnes , Federica Barontini , Filip Binkiewicz , Peter Black , Sebastian Bodenstedt , Leonardo Borgioli , Nikola Budjak , Benjamin Calmé , Fabio Carrillo , Nicola Cavalcanti , Changwei Chen , Haoxin Chen , Sihang Chen , Qihan Chen , Zhongyu Chen , Ziyang Chen , Shing Shin Cheng , Meiqing Cheng , Min Cheng , Zih-Yun Sarah Chiu , Xiangyu Chu , Camilo Correa-Gallego , Giulio Dagnino , Anton Deguet , Jacob Delgado , Jonathan C. DeLong , Kaizhong Deng , Alexander Dimitrakakis , Qingpeng Ding , Hao Ding , Giovanni Distefano , Daniel Donoho , Anqing Duan , Marco Esposito , Shane Farritor , Jad Fayad , Zahi Fayad , Mario Ferradosa , Filippo Filicori , Chelsea Finn , Philipp Fürnstahl , Jiawei Ge , Stamatia Giannarou , Xavier Giralt Ludevid , Frederic Giraud , Aditya Amit Godbole , Ken Goldberg , Antony Goldenberg , Diego Granero Marana , Xiaoqing Guo , Tamás Haidegger , Evan Hailey , Pascal Hansen , Ziyi Hao , Kush Hari , Kengo Hayashi , Jonathon Hawkins , Shelby Haworth , Ortrun Hellig , S. Duke Herrell , Zhouyang Hong , Andrew Howe , Junlei Hu , Zhaoyang Jacopo Hu , Ria Jain , Mohammad Rafiee Javazm , Howard Ji , Rui Ji , Jianmin Ji , Zhongliang Jiang , Dominic Jones , Jeffrey Jopling , Britton Jordan , Ran Ju , Michael Kam , Luoyao Kang , Fausto Kang , Siddhartha Kapuria , Peter Kazanzides , Sonika Kiehler , Ethan Kilmer , Ji Woong Kim , Przemysław Korzeniowski , Chandra Kuchi , Nithesh Kumar , Alan Kuntz , Federico Lavagno , Yu Chung Lee , Hao-Chih Lee , Hang Li , Zhen Li , Xiao Liang , Xinxin Lin , Jinsong Lin , Chang Liu , Fei Liu , Pei Liu , Yun-hui Liu , Wanli Liuchen , Eszter Lukács , Sareena Mann , Miles Mannas , Brett Marinelli , Sabina Martyniak , Francesco Marzola , Lorenzo Mazza , Xueyan Mei , Maria Clara Morais , Luigi Muratore , Chetan Reddy Narayanaswamy , Michał Naskręt , David Navarro-Alarcon , Cyrus Neary , Chi Kit Ng , Christopher Nguan , David Noonan , Ki Hwan Oh , Tom Christian Olesch , Allison M. Okamura , Justin Opfermann , Matteo Pescio , Doan Xuan Viet Pham , Tito Porras , Hongliang Ren , Ariel Rodriguez Jimenez , Ferdinando Rodriguez y Baena , Septimiu E. Salcudean , Asmitha Sathya , Preethi Satish , Lalithkumar Seenivasan , Jiaqi Shao , Yiqing Shen , Yu Sheng , Lucy XiaoYang Shi , Zoe Soulé , Stefanie Speidel , Mingwu Su , Jianhao Su , Idris Sunmola , Kristóf Takács , Yunxi Tang , Patrick Thornycroft , Yu Tian , Jordan Thompson , Mehmet K. Turkcan , Mathias Unberath , Pietro Valdastri , Carlos Vives , Quan Vuong , Martin Wagner , Farong Wang , Wei Wang , Lidian Wang , Chung-Pang Wang , Guankun Wang , Junyi Wang , Erqi Wang , Ziyi Wang , Tanner Watts , Wolfgang Wein , Yimeng Wu , Zijian Wu , Hongjun Wu , Luohong Wu , Jie Ying Wu , Junlin Wu , Victoria Wu , Kaixuan Wu , Mateusz Wójcikowski , Yunye Xiao , Nan Xiao , Wenxuan Xie , Hao Yang , Tianqi Yang , Yinuo Yang , Menglong Ye , Ryan S. Yeung , Nural Yilmaz , Chim Ho Yin , Michael Yip , Rayan Younis , Chenhao Yu , Sayem Nazmuz Zaman , Milos Zefran , Han Zhang , Yuelin Zhang , Yidong Zhang , Yanyong Zhang , Xuyang Zhang , Yameng Zhang , Joyce Zhang , Ning Zhong , Peng Zhou , Haoying Zhou , Xiuli Zuo , Nassir Navab , Mahdi Azizian , Sean D. Huver , Axel Krieger

We present a conceptual framework for training Vision-Language Models (VLMs) to perform Visual Perspective Taking (VPT), a core capability for embodied cognition essential for Human-Robot Interaction (HRI). As a first step toward this goal,…

Artificial Intelligence · Computer Science 2025-05-21 Joel Currie , Gioele Migno , Enrico Piacenti , Maria Elena Giannaccini , Patric Bach , Davide De Tommaso , Agnieszka Wykowska

Understanding and recognizing human-object interaction (HOI) is a pivotal application in AR/VR and robotics. Recent open-vocabulary HOI detection approaches depend exclusively on large language models for richer textual prompts, neglecting…

Computer Vision and Pattern Recognition · Computer Science 2025-08-18 Zhenhao Zhang , Hanqing Wang , Xiangyu Zeng , Ziyu Cheng , Jiaxin Liu , Haoyu Yan , Zhirui Liu , Kaiyang Ji , Tianxiang Gui , Ke Hu , Kangyi Chen , Yahao Fan , Mokai Pan

Embodied decision-making enables agents to translate high-level goals into executable actions through continuous interactions within the physical world, forming a cornerstone of general-purpose embodied intelligence. Large language models…

Artificial Intelligence · Computer Science 2025-10-15 Zixing Lei , Sheng Yin , Yichen Xiong , Yuanzhuo Ding , Wenhao Huang , Yuxi Wei , Qingyao Xu , Yiming Li , Weixin Li , Yunhong Wang , Siheng Chen

World models are emerging as a foundational paradigm for scalable, data-efficient embodied AI. In this work, we present GigaWorld-0, a unified world model framework designed explicitly as a data engine for Vision-Language-Action (VLA)…

This paper introduces CognitiveDog, a pioneering development of quadruped robot with Large Multi-modal Model (LMM) that is capable of not only communicating with humans verbally but also physically interacting with the environment through…

In this paper, we introduce a model of evolution and learning in robots that co-optimizes a distribution of latent design vectors (genotypes) and a mixture of control experts (neural modules), which are gated by the latent coordinates of…

Robotics · Computer Science 2026-05-26 Yibin Wang , Muhan Li , Zihan Guo , Sam Kriegman

Effective human-robot interaction requires emotionally rich multimodal expressions, yet most humanoid robots lack coordinated speech, facial expressions, and gestures. Meanwhile, real-world deployment demands on-device solutions that can…

Robotics · Computer Science 2026-02-10 Songhua Yang , Xuetao Li , Xuanye Fei , Mengde Li , Miao Li

Robotic manipulation has increasingly adopted vision-language-action (VLA) models, which achieve strong performance but typically require task-specific demonstrations and fine-tuning, and often generalize poorly under domain shift. We…

Robotics · Computer Science 2026-01-29 Brian Y. Tsui , Alan Y. Fang , Tiffany J. Hwu

Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and natural language goals. While recent vision-language models (VLMs) excel at static perception tasks, they struggle with the…

Artificial Intelligence · Computer Science 2025-07-15 Di Wu , Jiaxin Fan , Junzhe Zang , Guanbo Wang , Wei Yin , Wenhao Li , Bo Jin

Embodied reasoning systems integrate robotic hardware and cognitive processes to perform complex tasks, typically in response to a natural language query about a specific physical environment. This usually involves changing the belief about…

Vision-Language-Action Models (VLAs) have shown remarkable progress towards embodied intelligence. While their architecture partially resembles that of Large Language Models (LLMs), VLAs exhibit higher complexity due to their multi-modal…

Robotics · Computer Science 2026-03-06 Hugo Buurmeijer , Carmen Amo Alonso , Aiden Swann , Marco Pavone

Unified multimodal models can both understand and generate visual content within a single architecture. Existing models, however, remain data-hungry and too heavy for deployment on edge devices. We present Mobile-O, a compact…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Abdelrahman Shaker , Ahmed Heakl , Jaseel Muhammad , Ritesh Thawkar , Omkar Thawakar , Senmao Li , Hisham Cholakkal , Ian Reid , Eric P. Xing , Salman Khan , Fahad Shahbaz Khan

Recent large vision-language models (VLMs) remain fundamentally constrained by a persistent dichotomy: understanding and generation are treated as distinct problems, leading to fragmented architectures, cascaded pipelines, and misaligned…

Robot foundation models, particularly Vision-Language-Action (VLA) models, have garnered significant attention for their ability to enhance robot policy learning, greatly improving robot's generalization and robustness. OpenAI's recent…

Large Language Models have demonstrated remarkable reasoning capability in complex textual tasks. However, multimodal reasoning, which requires integrating visual and textual information, remains a significant challenge. Existing…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Yi Yang , Xiaoxuan He , Hongkun Pan , Xiyan Jiang , Yan Deng , Xingtao Yang , Haoyu Lu , Dacheng Yin , Fengyun Rao , Minfeng Zhu , Bo Zhang , Wei Chen

Embodied AI systems, including robots and autonomous vehicles, are increasingly integrated into real-world applications, where they encounter a range of vulnerabilities stemming from both environmental and system-level factors. These…

Cryptography and Security · Computer Science 2025-02-26 Wenpeng Xing , Minghao Li , Mohan Li , Meng Han

Long-horizon video-audio reasoning and fine-grained pixel understanding impose conflicting requirements on omnimodal models: dense temporal coverage demands many low-resolution frames, whereas precise grounding calls for high-resolution…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Hao Zhong , Muzhi Zhu , Zongze Du , Zheng Huang , Canyu Zhao , Mingyu Liu , Wen Wang , Hao Chen , Chunhua Shen

Bridging the gap between embodied intelligence and embedded deployment remains a key challenge in intelligent robotic systems, where perception, reasoning, and planning must operate under strict constraints on computation, memory, energy,…

Robotics · Computer Science 2026-05-19 Kuan Xu , Ruimeng Liu , Yizhuo Yang , Denan Liang , Tongxing Jin , Shenghai Yuan , Chen Wang , Lihua Xie
‹ Prev 1 8 9 10 Next ›