English
Related papers

Related papers: RoboMIND 2.0: A Multimodal, Bimanual Mobile Manipu…

200 papers

Humanoid robots, as general-purpose physical agents, must integrate both intelligent control and adaptive morphology to operate effectively in diverse real-world environments. While recent research has focused primarily on optimizing…

Robotics · Computer Science 2025-10-06 Guiliang Liu , Bo Yue , Yi Jin Kim , Kui Jia

Embodied foundation models are increasingly performant in real-world domains such as robotics or autonomous driving. These models are often deployed in interactive or assistive settings, where it is important that these assistive models…

Robotics · Computer Science 2026-03-06 Pradyumna Tambwekar , Andrew Silva , Deepak Gopinath , Jonathan DeCastro , Xiongyi Cui , Guy Rosman

We introduce HY-Embodied-0.5, a family of foundation models specifically designed for real-world embodied agents. To bridge the gap between general Vision-Language Models (VLMs) and the demands of embodied agents, our models are developed…

As robots increasingly enter human-centered environments, they must not only be able to navigate safely around humans, but also adhere to complex social norms. Humans often rely on non-verbal communication through gestures and facial…

Enhancing the spatial perception capabilities of mobile robots is crucial for achieving embodied Vision-and-Language Navigation (VLN). Although significant progress has been made in simulated environments, directly transferring these…

Artificial Intelligence · Computer Science 2026-03-03 Qianqian Bai , Zhongpu Chen , Ling Luo , Huaming Du , Yuqian Lei , Ziyun Jiao

Despite recent progress in Multi-Modal Large Language Models (MLLMs), it remains challenging to integrate diverse tasks ranging from pixel-level perception to high-fidelity generation. Existing approaches often suffer from either restricted…

Computation and Language · Computer Science 2026-01-29 Bin Zhu , Munan Ning , Peng Jin , Bin Lin , Jinfa Huang , Qi Song , Junwu Zhang , Zhenyu Tang , Mingjun Pan , Li Yuan

Recent advances in large-scale video world models have enabled increasingly realistic future prediction, raising the prospect of using generated videos as scalable supervision for robot learning. However, for embodied manipulation,…

Can we enable humanoid robots to generate rich, diverse, and expressive motions in the real world? We propose to learn a whole-body control policy on a human-sized robot to mimic human motions as realistic as possible. To train such a…

Robotics · Computer Science 2024-03-07 Xuxin Cheng , Yandong Ji , Junming Chen , Ruihan Yang , Ge Yang , Xiaolong Wang

Learning to execute long-horizon mobile manipulation tasks is crucial for advancing robotics in household and workplace settings. However, current approaches are typically data-inefficient, underscoring the need for improved models that…

Open-vocabulary mobile manipulation (OVMM) requires robots to follow language instructions, navigate, and manipulate while updating their world representation under dynamic environmental changes. However, most prior approaches update their…

Robotics · Computer Science 2026-04-15 Seongwon Cho , Daechul Ahn , Donghyun Shin , Hyeonbeom Choi , San Kim , Jonghyun Choi

We introduce GE-Sim 2.0 (Genie Envisioner World Simulator 2.0), a closed-loop video world simulator for robotic manipulation. Building on the action-conditioned video generation framework of Genie Envisioner, GE-Sim 2.0 is re-trained on…

Multimodal task specification is essential for enhanced robotic performance, where \textit{Cross-modality Alignment} enables the robot to holistically understand complex task instructions. Directly annotating multimodal instructions for…

Humans naturally process real-world multimodal information in a full-duplex manner. In artificial intelligence, replicating this capability is essential for advancing model development and deployment, particularly in embodied contexts. The…

Artificial Intelligence · Computer Science 2025-06-03 Yiqun Yao , Xiang Li , Xin Jiang , Xuezhi Fang , Naitong Yu , Aixin Sun , Yequan Wang

Human videos offer a scalable way to train robot manipulation policies, but lack the action labels needed by standard imitation learning algorithms. Existing cross-embodiment approaches try to map human motion to robot actions, but often…

Human behaviors in the real world naturally encode rich, long-term contextual information that can be leveraged to train embodied agents for perception, understanding, and acting. However, existing capture systems typically rely on costly…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Wenjia Wang , Liang Pan , Huaijin Pi , Yuke Lou , Xuqian Ren , Yifan Wu , Zhouyingcheng Liao , Lei Yang , Rishabh Dabral , Christian Theobalt , Taku Komura

Availability of large and diverse medical datasets is often challenged by privacy and data sharing restrictions. For successful application of machine learning techniques for disease diagnosis, prognosis, and precision medicine, large…

Egocentric human videos provide scalable demonstrations for imitation learning, but existing corpora often lack either fine-grained, temporally localized action descriptions or dexterous hand annotations. We introduce OpenEgo, a multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-09 Ahad Jawaid , Yu Xiang

Embodied learning for object-centric robotic manipulation is a rapidly developing and challenging area in embodied AI. It is crucial for advancing next-generation intelligent robots and has garnered significant interest recently. Unlike…

Robotics · Computer Science 2025-01-15 Ying Zheng , Lei Yao , Yuejiao Su , Yi Zhang , Yi Wang , Sicheng Zhao , Yiyi Zhang , Lap-Pui Chau

This paper investigates humanoid whole-body dexterous manipulation, where the efficient collection of high-quality demonstration data remains a central bottleneck. Existing teleoperation systems often suffer from limited portability,…

Robotics · Computer Science 2026-03-16 Liang Heng , Yihe Tang , Jiajun Xu , Henghui Bao , Di Huang , Yue Wang

Service robots in public spaces require real-time understanding of human behavioral intentions for natural interaction. We present a practical multimodal framework for frame-accurate human-robot interaction intent detection that fuses…

Robotics · Computer Science 2025-12-23 Farida Mohsen , Ali Safa
‹ Prev 1 8 9 10 Next ›