中文
相关论文

相关论文: Surfer: Progressive Reasoning with World Models fo…

200 篇论文

We introduce a method for real-time navigation and tracking with differentiably rendered world models. Learning models for control has led to impressive results in robotics and computer games, but this success has yet to be extended to…

机器学习 · 计算机科学 2022-01-26 Baris Kayalibay , Atanas Mirchev , Patrick van der Smagt , Justin Bayer

Human-robot interaction requires robots to process language incrementally, adapting their actions in real-time based on evolving speech input. Existing approaches to language-guided robot motion planning typically assume fully specified…

机器人学 · 计算机科学 2026-02-16 Mitchell Abrams , Thies Oelerich , Christian Hartl-Nesic , Andreas Kugi , Matthias Scheutz

A major challenge to deploying robots widely is navigation in human-populated environments, commonly referred to as social robot navigation. While the field of social navigation has advanced tremendously in recent years, the fair evaluation…

Visual navigation by mobile robots is classically tackled through SLAM plus optimal planning, and more recently through end-to-end training of policies implemented as deep networks. While the former are often limited to waypoint planning,…

人工智能 · 计算机科学 2021-11-30 Assem Sadek , Guillaume Bono , Boris Chidlovskii , Christian Wolf

An embodied system must not only model the patterns of the external world but also understand its own motion dynamics. A motion dynamic model is essential for efficient skill acquisition and effective planning. In this work, we introduce…

机器学习 · 计算机科学 2025-04-10 Chenjie Hao , Weyl Lu , Yifan Xu , Yubei Chen

The emerging integration of robots into everyday life brings several major challenges. Compared to classical industrial applications, more flexibility is needed in combination with real-time reactivity. Learning-based methods can train…

机器人学 · 计算机科学 2026-02-18 Thies Oelerich , Gerald Ebmer , Christian Hartl-Nesic , Andreas Kugi

The language-conditioned robotic manipulation aims to transfer natural language instructions into executable actions, from simple pick-and-place to tasks requiring intent recognition and visual reasoning. Inspired by the dual process theory…

With robotics rapidly advancing, more effective human-robot interaction is increasingly needed to realize the full potential of robots for society. While spoken language must be part of the solution, our ability to provide spoken language…

机器人学 · 计算机科学 2024-03-25 Matthew Marge , Carol Espy-Wilson , Nigel Ward

Model-based reinforcement learning (MBRL) has achieved remarkable success in robotics due to its high sample efficiency and planning capability. However, extending MBRL to physical multi-robot cooperation remains challenging due to the…

机器人学 · 计算机科学 2026-04-07 Zijie Zhao , Honglei Guo , Shengqian Chen , Kaixuan Xu , Bo Jiang , Yuanheng Zhu , Dongbin Zhao

Robotic manipulation in high-precision tasks is essential for numerous industrial and real-world applications where accuracy and speed are required. Yet current diffusion-based policy learning methods generally suffer from low computational…

机器人学 · 计算机科学 2025-06-23 Sen Wang , Le Wang , Sanping Zhou , Jingyi Tian , Jiayi Li , Haowen Sun , Wei Tang

For AI agents to be helpful to humans, they should be able to follow natural language instructions to complete everyday cooperative tasks in human environments. However, real human instructions inherently possess ambiguity, because the…

人工智能 · 计算机科学 2024-09-27 Yanming Wan , Yue Wu , Yiping Wang , Jiayuan Mao , Natasha Jaques

Recent advances in deep thinking models have demonstrated remarkable reasoning capabilities on mathematical and coding tasks. However, their effectiveness in embodied domains which require continuous interaction with environments through…

The automatic control of mobile devices is essential for efficiently performing complex tasks that involve multiple sequential steps. However, these tasks pose significant challenges due to the limited environmental information available at…

计算机视觉与模式识别 · 计算机科学 2025-09-08 Xiaoran Yin , Xu Luo , Hao Wu , Lianli Gao , Jingkuan Song

Recent Large Multimodal Models have demonstrated remarkable reasoning capabilities, especially in solving complex mathematical problems and realizing accurate spatial perception. Our key insight is that these emerging abilities can…

人工智能 · 计算机科学 2025-05-20 Weiliang Tang , Dong Jing , Jia-Hui Pan , Zhiwu Lu , Yun-Hui Liu , Li Erran Li , Mingyu Ding , Chi-Wing Fu

Robotic manipulation requires anticipating how the environment evolves in response to actions, yet most existing systems lack this predictive capability, often resulting in errors and inefficiency. While Vision-Language Models (VLMs)…

机器人学 · 计算机科学 2026-02-12 Songen Gu , Yunuo Cai , Tianyu Wang , Simo Wu , Yanwei Fu

Robotic assistance in scientific laboratories requires procedurally correct long-horizon manipulation, reliable execution under limited supervision, and robustness in low-demonstration regimes. Such conditions greatly challenge end-to-end…

机器人学 · 计算机科学 2026-02-11 Jinghan Yang , Jingyi Hou , Xinbo Yu , Wei He , Yifan Wu

Navigation guided by natural language instructions presents a challenging reasoning problem for instruction followers. Natural language instructions typically identify only a few high-level decisions and landmarks rather than complete…

Recent advances in video generation have shown promise for generating future scenarios, critical for planning and control in autonomous driving and embodied intelligence. However, real-world applications demand more than visually plausible…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Tianshuo Xu , Zhifei Chen , Leyi Wu , Hao Lu , Yuying Chen , Lihui Jiang , Bingbing Liu , Yingcong Chen

Embodied intelligence has witnessed remarkable progress in recent years, driven by advances in computer vision, natural language processing, and the rise of large-scale multimodal models. Among its core challenges, robot manipulation stands…

Vision-language models (VLMs) have exhibited impressive capabilities across diverse image understanding tasks, but still struggle in settings that require reasoning over extended sequences of camera frames from a video. This limits their…

计算与语言 · 计算机科学 2025-12-01 Philip Schroeder , Ondrej Biza , Thomas Weng , Hongyin Luo , James Glass