中文
相关论文

相关论文: EquAct: An SE(3)-Equivariant Multi-Task Transforme…

200 篇论文

A truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Jia-Xing Zhong , Ta-Ying Cheng , Yuhang He , Kai Lu , Kaichen Zhou , Andrew Markham , Niki Trigoni

Transformers have revolutionized vision and natural language processing with their ability to scale with large datasets. But in robotic manipulation, data is both limited and expensive. Can manipulation still benefit from Transformers with…

机器人学 · 计算机科学 2022-11-14 Mohit Shridhar , Lucas Manuelli , Dieter Fox

This work introduces E3x, a software package for building neural networks that are equivariant with respect to the Euclidean group $\mathrm{E}(3)$, consisting of translations, rotations, and reflections of three-dimensional space. Compared…

机器学习 · 计算机科学 2024-11-12 Oliver T. Unke , Hartmut Maennel

Humans perceive and interact with the world with the awareness of equivariance, facilitating us in manipulating different objects in diverse poses. For robotic manipulation, such equivariance also exists in many scenarios. For example, no…

机器人学 · 计算机科学 2024-08-08 Yue Chen , Chenrui Tie , Ruihai Wu , Hao Dong

Accurately predicting 3D structures and dynamics of physical systems is crucial in scientific applications. Existing approaches that rely on geometric Graph Neural Networks (GNNs) effectively enforce $\mathrm{E}(3)$-equivariance, but they…

机器学习 · 计算机科学 2025-02-20 Zongzhao Li , Jiacheng Cen , Bing Su , Wenbing Huang , Tingyang Xu , Yu Rong , Deli Zhao

Vision-language-action policies learn manipulation skills across tasks, environments and embodiments through large-scale pre-training. However, their ability to generalize to novel robot configurations remains limited. Most approaches…

机器人学 · 计算机科学 2025-09-19 Anzhe Chen , Yifei Yang , Zhenjie Zhu , Kechun Xu , Zhongxiang Zhou , Rong Xiong , Yue Wang

Diffusion Policies are effective at learning closed-loop manipulation policies from human demonstrations but generalize poorly to novel arrangements of objects in 3D space, hurting real-world performance. To address this issue, we propose…

机器人学 · 计算机科学 2025-07-03 Xupeng Zhu , Fan Wang , Robin Walters , Jane Shi

Recently, equivariant neural network models have been shown to improve sample efficiency for tasks in computer vision and reinforcement learning. This paper explores this idea in the context of on-robot policy learning in which a policy…

机器人学 · 计算机科学 2022-10-19 Dian Wang , Mingxi Jia , Xupeng Zhu , Robin Walters , Robert Platt

Visuotactile policy learning augments vision-only policies with tactile input, facilitating contact-rich manipulation. However, the high cost of tactile data collection makes sample efficiency the key requirement for developing visuotactile…

机器人学 · 计算机科学 2025-11-12 Yizhe Zhu , Zhang Ye , Boce Hu , Haibo Zhao , Yu Qi , Dian Wang , Robert Platt

Contemporary autoregressive transformers operate in open loop: each hidden state is computed in a single forward pass and never revised, causing errors to propagate uncorrected through the sequence. We identify this open-loop bottleneck as…

机器学习 · 计算机科学 2025-12-01 Akbar Anbar Jafari , Gholamreza Anbarjafari

We propose a method for 3D shape reconstruction from unoriented point clouds. Our method consists of a novel SE(3)-equivariant coordinate-based network (TF-ONet), that parametrizes the occupancy field of the shape and respects the inherent…

计算机视觉与模式识别 · 计算机科学 2023-02-13 Evangelos Chatzipantazis , Stefanos Pertigkiozoglou , Edgar Dobriban , Kostas Daniilidis

Shape assembly aims to reassemble parts (or fragments) into a complete object, which is a common task in our daily life. Different from the semantic part assembly (e.g., assembling a chair's semantic parts like legs into a whole chair),…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Ruihai Wu , Chenrui Tie , Yushi Du , Yan Zhao , Hao Dong

We present RiEMann, an end-to-end near Real-time SE(3)-Equivariant Robot Manipulation imitation learning framework from scene point cloud input. Compared to previous methods that rely on descriptor field matching, RiEMann directly predicts…

机器人学 · 计算机科学 2024-10-04 Chongkai Gao , Zhengrong Xue , Shuying Deng , Tianhai Liang , Siqi Yang , Lin Shao , Huazhe Xu

Learning to predict agent motions with relationship reasoning is important for many applications. In motion prediction tasks, maintaining motion equivariance under Euclidean geometric transformations and invariance of agent interaction is a…

计算机视觉与模式识别 · 计算机科学 2023-03-28 Chenxin Xu , Robby T. Tan , Yuhong Tan , Siheng Chen , Yu Guang Wang , Xinchao Wang , Yanfeng Wang

Imitation learning, e.g., diffusion policy, has been proven effective in various robotic manipulation tasks. However, extensive demonstrations are required for policy robustness and generalization. To reduce the demonstration reliance, we…

机器人学 · 计算机科学 2025-03-04 Chenrui Tie , Yue Chen , Ruihai Wu , Boxuan Dong , Zeyi Li , Chongkai Gao , Hao Dong

3D perceptual representations are well suited for robot manipulation as they easily encode occlusions and simplify spatial reasoning. Many manipulation tasks require high spatial precision in end-effector pose prediction, which typically…

机器人学 · 计算机科学 2023-10-23 Theophile Gervet , Zhou Xian , Nikolaos Gkanatsios , Katerina Fragkiadaki

Visual Imitation learning has achieved remarkable progress in robotic manipulation, yet generalization to unseen objects, scene layouts, and camera viewpoints remains a key challenge. Recent advances address this by using 3D point clouds,…

机器人学 · 计算机科学 2025-11-11 Zhiyuan Zhang , Zhengtong Xu , Jai Nanda Lakamsani , Yu She

Spatial understanding is a critical aspect of most robotic tasks, particularly when generalization is important. Despite the impressive results of deep generative models in complex manipulation tasks, the absence of a representation that…

机器人学 · 计算机科学 2024-09-10 Niklas Funk , Julen Urain , Joao Carvalho , Vignesh Prasad , Georgia Chalvatzaki , Jan Peters

Robotic manipulation in unstructured environments requires the generation of robust and long-horizon trajectory-level policy with conditions of perceptual observations and benefits from the advantages of SE(3)-equivariant diffusion models…

机器人学 · 计算机科学 2025-09-30 Zhitao Wang , Yanke Wang , Jiangtao Wen , Roberto Horowitz , Yuxing Han

We present e3nn, a generalized framework for creating E(3) equivariant trainable functions, also known as Euclidean neural networks. e3nn naturally operates on geometry and geometric tensors that describe systems in 3D and transform…

机器学习 · 计算机科学 2022-07-21 Mario Geiger , Tess Smidt