English
Related papers

Related papers: EquAct: An SE(3)-Equivariant Multi-Task Transforme…

200 papers

We present DM$^3$-Nav, a fully decentralized multi-agent semantic navigation system supporting multimodal open-vocabulary goal specification and multi-object missions. In our setting, decentralization implies operation without a central…

Multiagent Systems · Computer Science 2026-04-27 Amin Kashiri , Atharva Jamsandekar , Yasin Yazıcıoğlu

Aerospace embodied intelligence aims to empower unmanned aerial vehicles (UAVs) and other aerospace platforms to achieve autonomous perception, cognition, and action, as well as egocentric active interaction with humans and the environment.…

Robotics · Computer Science 2025-11-24 Fanglong Yao , Yuanchang Yue , Youzhi Liu , Xian Sun , Kun Fu

Language-guided long-horizon mobile manipulation has long been a grand challenge in embodied semantic reasoning, generalizable manipulation, and adaptive locomotion. Three fundamental limitations hinder progress: First, although large…

Robotics · Computer Science 2025-08-12 Kaijun Wang , Liqin Lu , Mingyu Liu , Jianuo Jiang , Zeju Li , Bolin Zhang , Wancai Zheng , Xinyi Yu , Hao Chen , Chunhua Shen

Recent years, multimodal models have made remarkable strides and pave the way for intelligent browser use agents. However, when solving tasks on real world webpages in multi-turn, long-horizon trajectories, current agents still suffer from…

Artificial Intelligence · Computer Science 2025-09-26 Kaiwen He , Zhiwei Wang , Chenyi Zhuang , Jinjie Gu

For 3D object manipulation, methods that build an explicit 3D representation perform better than those relying only on camera images. But using explicit 3D representations like voxels comes at large computing cost, adversely affecting…

Robotics · Computer Science 2023-06-27 Ankit Goyal , Jie Xu , Yijie Guo , Valts Blukis , Yu-Wei Chao , Dieter Fox

Recent advances in imitation learning for 3D robotic manipulation have shown promising results with diffusion-based policies. However, achieving human-level dexterity requires seamless integration of geometric precision and semantic…

Learning for robot navigation presents a critical and challenging task. The scarcity and costliness of real-world datasets necessitate efficient learning approaches. In this letter, we exploit Euclidean symmetry in planning for 2D…

Robotics · Computer Science 2024-01-30 Linfeng Zhao , Hongyu Li , Taskin Padir , Huaizu Jiang , Lawson L. S. Wong

Integrating a notion of symmetry into point cloud neural networks is a provably effective way to improve their generalization capability. Of particular interest are $E(3)$ equivariant point cloud networks where Euclidean transformations…

Machine Learning · Computer Science 2024-02-14 Matan Atzmon , Jiahui Huang , Francis Williams , Or Litany

Generalizing language-conditioned robotic policies to new tasks remains a significant challenge, hampered by the lack of suitable simulation benchmarks. In this paper, we address this gap by introducing GemBench, a novel benchmark to assess…

Robotics · Computer Science 2025-03-04 Ricardo Garcia , Shizhe Chen , Cordelia Schmid

To enable robots to comprehend high-level human instructions and perform complex tasks, a key challenge lies in achieving comprehensive scene understanding: interpreting and interacting with the 3D environment in a meaningful way. This…

360 cameras capture the entire surrounding environment with a large FoV, exhibiting comprehensive visual information to directly infer the 3D structures, e.g., depth and surface normal, and semantic information simultaneously. Existing…

Computer Vision and Pattern Recognition · Computer Science 2024-08-20 Hao Ai , Lin Wang

We present ReAct!, an interactive tool for high-level reasoning for cognitive robotic applications. ReAct! enables robotic researchers to describe robots' actions and change in dynamic domains, without having to know about the syntactic and…

Artificial Intelligence · Computer Science 2026-05-14 Zeynep Dogmus , Esra Erdem , Volkan Patoglu

Recent advanced vision-language models(VLMs) have demonstrated strong performance on passive, offline image and video understanding tasks. However, their effectiveness in embodied settings, which require online interaction and active scene…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Mingxian Lin , Wei Huang , Yitang Li , Chengjie Jiang , Kui Wu , Fangwei Zhong , Shengju Qian , Xin Wang , Xiaojuan Qi

Popular representation learning methods encourage feature invariance under transformations applied at the input. However, in 3D perception tasks like object localization and segmentation, outputs are naturally equivariant to some…

Computer Vision and Pattern Recognition · Computer Science 2024-04-19 Deepti Hegde , Suhas Lohit , Kuan-Chuan Peng , Michael J. Jones , Vishal M. Patel

Analyzing volumetric data with rotational invariance or equivariance is an active topic in current research. Existing deep-learning approaches utilize either group convolutional networks limited to discrete rotations or steerable…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Dmitrii Zhemchuzhnikov , Sergei Grudinin

Equivariant models have recently been shown to improve the data efficiency of diffusion policy by a significant margin. However, prior work that explored this direction focused primarily on point cloud inputs generated by multiple cameras…

Robotics · Computer Science 2025-10-31 Boce Hu , Dian Wang , David Klee , Heng Tian , Xupeng Zhu , Haojie Huang , Robert Platt , Robin Walters

Humanoid robots promise to operate in everyday human environments without requiring modifications to the surroundings. Among the many skills needed, opening doors is essential, as doors are the most common gateways in built spaces and often…

Equivariance has been a long-standing concern in various fields ranging from computer vision to physical modeling. Most previous methods struggle with generality, simplicity, and expressiveness -- some are designed ad hoc for specific data…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Shitong Luo , Jiahan Li , Jiaqi Guan , Yufeng Su , Chaoran Cheng , Jian Peng , Jianzhu Ma

Natural language is one of the most intuitive ways to express human intent. However, translating instructions and commands towards robotic motion generation and deployment in the real world is far from being an easy task. The challenge of…

Robotics · Computer Science 2022-09-20 Arthur Bucker , Luis Figueredo , Sami Haddadin , Ashish Kapoor , Shuang Ma , Sai Vemprala , Rogerio Bonatti

This paper presents DNAct, a language-conditioned multi-task policy framework that integrates neural rendering pre-training and diffusion training to enforce multi-modality learning in action sequence spaces. To learn a generalizable…

Robotics · Computer Science 2024-03-11 Ge Yan , Yueh-Hua Wu , Xiaolong Wang