English
Related papers

Related papers: MP5: A Multi-modal Open-ended Embodied System in M…

200 papers

The ability to simulate the effects of future actions on the world is a crucial ability of intelligent embodied agents, enabling agents to anticipate the effects of their actions and make plans accordingly. While a large body of existing…

Computer Vision and Pattern Recognition · Computer Science 2025-05-12 Siyuan Zhou , Yilun Du , Yuncong Yang , Lei Han , Peihao Chen , Dit-Yan Yeung , Chuang Gan

Replicating human-level intelligence in the execution of embodied tasks remains challenging due to the unconstrained nature of real-world environments. Novel use of large language models (LLMs) for task planning seeks to address the…

Large language model (LLM) based agents have shown great potential in following human instructions and automatically completing various tasks. To complete a task, the agent needs to decompose it into easily executed steps by planning.…

Computation and Language · Computer Science 2025-06-02 Weihong Du , Wenrui Liao , Binyu Yan , Hongru Liang , Anthony G. Cohn , Wenqiang Lei

In recent years, as machine learning, particularly for vision and language understanding, has been improved, research in embedded AI has also evolved. VOYAGER is a well-known LLM-based embodied AI that enables autonomous exploration in the…

Artificial Intelligence · Computer Science 2024-06-05 Wakana Haijima , Kou Nakakubo , Masahiro Suzuki , Yutaka Matsuo

We introduce RynnEC, a video multimodal large language model designed for embodied cognition. Built upon a general-purpose vision-language foundation model, RynnEC incorporates a region encoder and a mask decoder, enabling flexible…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Ronghao Dang , Yuqian Yuan , Yunxuan Mao , Kehan Li , Jiangpin Liu , Zhikai Wang , Xin Li , Fan Wang , Deli Zhao

Human-robot interaction is increasingly moving toward multi-robot, socially grounded environments. Existing systems struggle to integrate multimodal perception, embodied expression, and coordinated decision-making in a unified framework.…

Robotics · Computer Science 2026-03-25 Shaid Hasan , Breenice Lee , Sujan Sarker , Tariq Iqbal

Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks, yet they face significant challenges in embodied task planning scenarios that require continuous environmental understanding and action generation.…

Computation and Language · Computer Science 2025-07-01 Zhaoye Fei , Li Ji , Siyin Wang , Junhao Shi , Jingjing Gong , Xipeng Qiu

Recent efforts on training visual navigation agents conditioned on language using deep reinforcement learning have been successful in learning policies for different multimodal tasks, such as semantic goal navigation and embodied question…

Machine Learning · Computer Science 2019-02-05 Devendra Singh Chaplot , Lisa Lee , Ruslan Salakhutdinov , Devi Parikh , Dhruv Batra

Active vision, also known as active perception, refers to the process of actively selecting where and how to look in order to gather task-relevant information. It is a critical component of efficient perception and decision-making in humans…

Computer Vision and Pattern Recognition · Computer Science 2025-05-28 Muzhi Zhu , Hao Zhong , Canyu Zhao , Zongze Du , Zheng Huang , Mingyu Liu , Hao Chen , Cheng Zou , Jingdong Chen , Ming Yang , Chunhua Shen

We present an embodied AI system which receives open-ended natural language instructions from a human, and controls two arms to collaboratively accomplish potentially long-horizon tasks over a large workspace. Our system is modular: it…

In recent years, multimodal large language models (MLLMs) have shown remarkable capabilities in tasks like visual question answering and common sense reasoning, while visual perception models have made significant strides in perception…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Guanqun Wang , Xinyu Wei , Jiaming Liu , Ray Zhang , Yichi Zhang , Kevin Zhang , Maurice Chong , Shanghang Zhang

Embodied Reference Understanding requires identifying a target object in a visual scene based on both language instructions and pointing cues. While prior works have shown progress in open-vocabulary object detection, they often fail in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-30 Fevziye Irem Eyiokur , Dogucan Yaman , Hazım Kemal Ekenel , Alexander Waibel

Emergent communication offers insight into how agents develop shared structured representations, yet most research assumes homogeneous modalities or aligned representational spaces, overlooking the perceptual heterogeneity of real-world…

Multiagent Systems · Computer Science 2026-01-30 Naomi Pitzer , Daniela Mihai

The human ability to easily solve multimodal tasks in context (i.e., with only a few demonstrations or simple instructions), is what current multimodal systems have largely struggled to imitate. In this work, we demonstrate that the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-09 Quan Sun , Yufeng Cui , Xiaosong Zhang , Fan Zhang , Qiying Yu , Zhengxiong Luo , Yueze Wang , Yongming Rao , Jingjing Liu , Tiejun Huang , Xinlong Wang

Human beings possess the capability to multiply a melange of multisensory cues while actively exploring and interacting with the 3D world. Current multi-modal large language models, however, passively absorb sensory data as inputs, lacking…

Computer Vision and Pattern Recognition · Computer Science 2024-01-17 Yining Hong , Zishuo Zheng , Peihao Chen , Yian Wang , Junyan Li , Chuang Gan

Recent advancements in 3D reconstruction and neural rendering have enhanced the creation of high-quality digital assets, yet existing methods struggle to generalize across varying object shapes, textures, and occlusions. While Next Best…

Robotics · Computer Science 2024-09-25 Zhenghao Qi , Shenghai Yuan , Fen Liu , Haozhi Cao , Tianchen Deng , Jianfei Yang , Lihua Xie

Creating AI systems that can interact with environments over long periods, similar to human cognition, has been a longstanding research goal. Recent advancements in multimodal large language models (MLLMs) have made significant strides in…

The progression to "Pervasive Augmented Reality" envisions easy access to multimodal information continuously. However, in many everyday scenarios, users are occupied physically, cognitively or socially. This may increase the friction to…

Human-Computer Interaction · Computer Science 2024-05-08 Jiahao Nick Li , Yan Xu , Tovi Grossman , Stephanie Santosa , Michelle Li

In long-horizon open-world multi-agent systems, existing methods often treat local anomalies as automatic triggers for communication. This default design introduces coordination noise, interrupts local execution, and overuses public…

Multiagent Systems · Computer Science 2026-04-22 HuaDong Jian , Chenghao Li , Haoyu Wang , Jiajia Shuai , Jinyu Guo , Yang Yang , Chaoning Zhang

In recent years, soft prompt learning methods have been proposed to fine-tune large-scale vision-language pre-trained models for various downstream tasks. These methods typically combine learnable textual tokens with class tokens as input…

Computer Vision and Pattern Recognition · Computer Science 2024-05-01 Yingjie Tian , Yiqi Wang , Xianda Guo , Zheng Zhu , Long Chen
‹ Prev 1 4 5 6 7 8 10 Next ›