English
Related papers

Related papers: MiMo-Embodied: X-Embodied Foundation Model Technic…

200 papers

Multimodal learning, especially large-scale multimodal pre-training, has developed rapidly over the past few years and led to the greatest advances in artificial intelligence (AI). Despite its effectiveness, understanding the underlying…

Neural and Evolutionary Computing · Computer Science 2022-08-18 Haoyu Lu , Qiongyi Zhou , Nanyi Fei , Zhiwu Lu , Mingyu Ding , Jingyuan Wen , Changde Du , Xin Zhao , Hao Sun , Huiguang He , Ji-Rong Wen

Learning universal policies from cross-embodied data remains a fundamental challenge in robotics. Although Vision-Language-Action (VLA) models are pre-trained on large and diverse datasets, they typically rely on embodiment-specific…

Robotics · Computer Science 2026-05-26 Boyu Li , Chaoyi Xu , Haoqi Yuan , Xinrun Xu , Börje F. Karlsson , Dongbin Zhao , Haoran Li , Zongqing Lu

Robotic foundation models trained on large-scale manipulation datasets have shown promise in learning generalist policies, but they often overfit to specific viewpoints, robot arms, and especially parallel-jaw grippers due to dataset…

Robotics · Computer Science 2026-01-15 Tong Wu , Shoujie Li , Junhao Gong , Changqing Guo , Xingting Li , Shilong Mu , Wenbo Ding

Embodied AI requires agents that perceive, act, and anticipate how actions reshape future world states. World models serve as internal simulators that capture environment dynamics, enabling forward and counterfactual rollouts to support…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Xinqing Li , Xin He , Le Zhang , Min Wu , Xiaoli Li , Yun Liu

Leveraging Multi-modal Large Language Models (MLLMs) to create embodied agents offers a promising avenue for tackling real-world tasks. While language-centric embodied agents have garnered substantial attention, MLLM-based embodied agents…

Vision-language navigation requires agents to reason and act under constraints of embodiment. While vision-language models (VLMs) demonstrate strong generalization, current benchmarks provide limited understanding of how embodiment -- i.e.,…

Robotics · Computer Science 2025-12-23 Tin Stribor Sohn , Maximilian Dillitzer , Jason J. Corso , Eric Sax

In this paper, we introduce MIO, a novel foundation model built on multimodal tokens, capable of understanding and generating speech, text, images, and videos in an end-to-end, autoregressive manner. While the emergence of large language…

Existing language-driven embodied navigation paradigms face challenges in functional buildings (FBs) with highly similar features, as they lack the ability to effectively utilize priori spatial knowledge. To tackle this issue, we propose a…

Robotics · Computer Science 2026-03-11 Jiang Gao , Xiangyu Dong , Haozhou Li , Haoran Zhao , Yaoming Zhou , Xiaoguang Ma

With the rapid advancement of low-altitude remote sensing and Vision-Language Models (VLMs), Embodied Agents based on Unmanned Aerial Vehicles (UAVs) have shown significant potential in autonomous tasks. However, current evaluation methods…

Robotics · Computer Science 2025-12-09 Mingning Guo , Mengwei Wu , Jiarun He , Shaoxian Li , Haifeng Li , Chao Tao

With the rapid penetration of artificial intelligence across industries and scenarios, a key challenge in building the next-generation intelligent core lies in effectively integrating the language understanding capabilities of foundation…

Artificial Intelligence · Computer Science 2025-06-03 Liang Geng

We present an embodied AI system which receives open-ended natural language instructions from a human, and controls two arms to collaboratively accomplish potentially long-horizon tasks over a large workspace. Our system is modular: it…

We present Emu, a Transformer-based multimodal foundation model, which can seamlessly generate images and texts in multimodal context. This omnivore model can take in any single-modality or multimodal data input indiscriminately (e.g.,…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Quan Sun , Qiying Yu , Yufeng Cui , Fan Zhang , Xiaosong Zhang , Yueze Wang , Hongcheng Gao , Jingjing Liu , Tiejun Huang , Xinlong Wang

Multimodal Large Language Models (MLLMs) have demonstrated extraordinary progress in bridging textual and visual inputs. However, MLLMs still face challenges in situated physical and social interactions in sensorally rich, multimodal and…

Neurons and Cognition · Quantitative Biology 2025-10-17 Akila Kadambi , Lisa Aziz-Zadeh , Antonio Damasio , Marco Iacoboni , Srini Narayanan

Multimodal large language models (MLLMs) have advanced vision-language reasoning and are increasingly deployed in embodied agents. However, significant limitations remain: MLLMs generalize poorly across digital-physical spaces and…

Embodied Artificial Intelligence (Embodied AI) integrates perception, cognition, planning, and interaction into agents that operate in open-world, safety-critical environments. As these systems gain autonomy and enter domains such as…

Embodied Vision-Language Models (VLMs) have demonstrated impressive performance and generalization in robotics, particularly within Vision-Language-Action frameworks. However, a significant gap remains between the high-level semantic focus…

Computer Vision and Pattern Recognition · Computer Science 2026-05-28 Ruowen Zhao , Bangguo Li , Zuyan Liu , Yinan Liang , Junliang Ye , Fangfu Liu , Diankun Wu , Zhengyi Wang , Xumin Yu , Yongming Rao , Han Hu , Jun Zhu

We introduce iFlyBot-VLM, a general-purpose Vision-Language Model (VLM) used to improve the domain of Embodied Intelligence. The central objective of iFlyBot-VLM is to bridge the cross-modal semantic gap between high-dimensional…

Robotics · Computer Science 2025-11-10 Xin Nie , Zhiyuan Cheng , Yuan Zhang , Chao Ji , Jiajia Wu , Yuhan Zhang , Jia Pan

Embodied Artificial Intelligence (AI) is an intelligent system paradigm for achieving Artificial General Intelligence (AGI), serving as the cornerstone for various applications and driving the evolution from cyberspace to physical systems.…

Artificial Intelligence · Computer Science 2025-09-25 Tongtong Feng , Xin Wang , Yu-Gang Jiang , Wenwu Zhu

Embodied planning requires agents to make coherent multi-step decisions based on dynamic visual observations and natural language goals. While recent vision-language models (VLMs) excel at static perception tasks, they struggle with the…

Artificial Intelligence · Computer Science 2025-07-15 Di Wu , Jiaxin Fan , Junzhe Zang , Guanbo Wang , Wei Yin , Wenhao Li , Bo Jin

Robots are traditionally bounded by a fixed embodiment during their operational lifetime, which limits their ability to adapt to their surroundings. Co-optimizing control and morphology of a robot, however, is often inefficient due to the…

Robotics · Computer Science 2022-12-20 Chen Yu , Weinan Zhang , Hang Lai , Zheng Tian , Laurent Kneip , Jun Wang
‹ Prev 1 3 4 5 6 7 10 Next ›