English
Related papers

Related papers: Being-0: A Humanoid Robotic Agent with Vision-Lang…

200 papers

Multimodal Large Language Models (MLLMs) are evolving from passive observers into active agents, solving problems through Visual Expansion (invoking visual tools) and Knowledge Expansion (open-web search). However, existing evaluations fall…

Artificial Intelligence · Computer Science 2026-04-06 Qianshan Wei , Yishan Yang , Siyi Wang , Jinglin Chen , Binyu Wang , Jiaming Wang , Shuang Chen , Zechen Li , Yang Shi , Yuqi Tang , Weining Wang , Yi Yu , Chaoyou Fu , Qi Li , Yi-Fan Zhang

Embodied agents operating in household environments must interpret ambiguous and under-specified human instructions. A capable household robot should recognize ambiguity and ask relevant clarification questions to infer the user intent…

Artificial Intelligence · Computer Science 2025-10-06 Ram Ramrakhya , Matthew Chang , Xavier Puig , Ruta Desai , Zsolt Kira , Roozbeh Mottaghi

Language-conditioned robotic skills make it possible to apply the high-level reasoning of Large Language Models (LLMs) to low-level robotic control. A remaining challenge is to acquire a diverse set of fundamental skills. Existing…

Robotics · Computer Science 2024-08-19 Xufeng Zhao , Cornelius Weber , Stefan Wermter

Interpreting human intent accurately is a central challenge in human-robot interaction (HRI) and a key requirement for achieving more natural and intuitive collaboration between humans and machines. This work presents a novel multimodal HRI…

Robotics · Computer Science 2026-02-25 Guanting Shen , Zi Tian

Deep Learning has revolutionized our ability to solve complex problems such as Vision-and-Language Navigation (VLN). This task requires the agent to navigate to a goal purely based on visual sensory inputs given natural language…

Robotics · Computer Science 2021-04-22 Muhammad Zubair Irshad , Chih-Yao Ma , Zsolt Kira

Hierarchical control for robotics has long been plagued by the need to have a well defined interface layer to communicate between high-level task planners and low-level policies. With the advent of LLMs, language has been emerging as a…

Robotics · Computer Science 2025-07-09 Yide Shentu , Philipp Wu , Aravind Rajeswaran , Pieter Abbeel

Facing increasingly complex BIM authoring software and the accompanying expensive learning costs, designers often seek to interact with the software in a more intelligent and lightweight manner. They aim to automate modeling workflows,…

Human-Computer Interaction · Computer Science 2024-06-26 Changyu Du , Stavros Nousias , André Borrmann

Current vision-language navigation methods face substantial bottlenecks regarding heterogeneous robot compatibility, real-time performance, and navigation safety. Furthermore, they struggle to support open-vocabulary semantic generalization…

Robotics · Computer Science 2026-04-06 Mingao Tan , Yiyang Li , Shanze Wang , Xinming Zhang , Wei Zhang

We introduce a novel framework for automatic behavior tree (BT) construction in heterogeneous multi-robot systems, designed to address the challenges of adaptability and robustness in dynamic environments. Traditional robots are limited by…

Robotics · Computer Science 2025-10-14 Chaoran Wang , Jingyuan Sun , Yanhui Zhang , Mingyu Zhang , Changju Wu

Humans act with context and intention, with reasoning playing a central role. While internet-scale data has enabled broad reasoning capabilities in AI systems, grounding these abilities in physical action remains a major challenge. We…

Robotics · Computer Science 2025-12-11 Peijun Tang , Shangjin Xie , Binyan Sun , Baifu Huang , Kuncheng Luo , Haotian Yang , Weiqi Jin , Jianan Wang

Interactive robot learning is a challenging problem as the robot is present with human users who expect the robot to learn novel skills to solve novel tasks perpetually with sample efficiency. In this work we present a framework for robots…

Robotics · Computer Science 2026-03-31 Weiwei Gu , Suresh Kondepudi , Anmol Gupta , Lixiao Huang , Nakul Gopalan

Large language models (LLMs) and vision-language models (VLMs) have the potential to transform biological research by enabling autonomous experimentation. Yet, their application remains constrained by rigid protocol design, limited…

Robotics · Computer Science 2025-07-03 Yibo Qiu , Zan Huang , Zhiyu Wang , Handi Liu , Yiling Qiao , Yifeng Hu , Shu'ang Sun , Hangke Peng , Ronald X Xu , Mingzhai Sun

Vision-language models (VLMs) have excelled in multimodal tasks, but adapting them to embodied decision-making in open-world environments presents challenges. One critical issue is bridging the gap between discrete entities in low-level…

Computer Vision and Pattern Recognition · Computer Science 2025-03-21 Shaofei Cai , Zihao Wang , Kewei Lian , Zhancun Mu , Xiaojian Ma , Anji Liu , Yitao Liang

Language is an effective medium for bi-directional communication in human-robot teams. To infer the meaning of many instructions, robots need to construct a model of their surroundings that describe the spatial, semantic, and metric…

Robotics · Computer Science 2019-09-24 Ethan Fahnestock , Siddharth Patki , Thomas M. Howard

In this work, we introduce HoloBrain-0, a comprehensive Vision-Language-Action (VLA) framework that bridges the gap between foundation model research and reliable real-world robot deployment. The core of our system is a novel VLA…

Existing AutoML systems have advanced the automation of machine learning (ML); however, they still require substantial manual configuration and expert input, particularly when handling multimodal data. We introduce MLZero, a novel…

Mobile manipulators are increasingly deployed in human-centered environments to perform tasks. While completing such tasks, they should also be able to communicate their intent to the people around them using expressive robot behaviors.…

Robotics · Computer Science 2026-04-24 Souren Pashangpour , Haitong Wang , Matthew Lisondra , Goldie Nejat

Multimodal Large Language Models have shown promising capabilities in bridging visual and textual reasoning, yet their reasoning capabilities in Open-Vocabulary Human-Object Interaction (OV-HOI) are limited by cross-modal hallucinations and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Zhenlong Yuan , Yue Wang , Dapeng Zhang , Kejin Cui , Rui Chen , Jing Tang , Lei Sun , Hongwei Yu , Chengxuan Qian , Xiangxiang Chu , Shuo Li , Yuyin Zhou

Building general-purpose embodied agents across diverse hardware remains a central challenge in robotics, often framed as the ''one-brain, many-forms'' paradigm. Progress is hindered by fragmented data, inconsistent representations, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-15 Yandan Yang , Shuang Zeng , Tong Lin , Xinyuan Chang , Dekang Qi , Junjin Xiao , Haoyun Liu , Ronghan Chen , Yuzhi Chen , Dongjie Huo , Feng Xiong , Xing Wei , Zhiheng Ma , Mu Xu

Language model (LM)-based agents have demonstrated promising capabilities in automating complex tasks from natural language instructions, yet they continue to struggle with long-horizon planning and reasoning. To address this, we propose an…

Artificial Intelligence · Computer Science 2026-05-05 Wenyi Wu , Sibo Zhu , Kun Zhou , Biwei Huang