中文
相关论文

相关论文: Canonical Policy: Learning Canonical 3D Representa…

200 篇论文

Existing approaches for transporting and manipulating cable-suspended loads using multiple UAVs along reference trajectories typically rely on either centralized control architectures or reliable inter-agent communication. In this work, we…

机器人学 · 计算机科学 2025-10-21 Shantnav Agarwal , Javier Alonso-Mora , Sihao Sun

Conventional and deep learning-based methods have shown great potential in the medical imaging domain, as means for deriving diagnostic, prognostic, and predictive biomarkers, and by contributing to precision medicine. However, these…

With the capacity of modeling long-range dependencies in sequential data, transformers have shown remarkable performances in a variety of generative tasks such as image, audio, and text generation. Yet, taming them in generating less…

计算机视觉与模式识别 · 计算机科学 2022-04-06 An-Chieh Cheng , Xueting Li , Sifei Liu , Min Sun , Ming-Hsuan Yang

Humans can imagine goal states during planning and perform actions to match those goals. In this work, we propose Imagination Policy, a novel multi-task key-frame policy network for solving high-precision pick and place tasks. Instead of…

In robot learning, the observation space is crucial due to the distinct characteristics of different modalities, which can potentially become a bottleneck alongside policy design. In this study, we explore the influence of various…

机器人学 · 计算机科学 2024-10-23 Haoyi Zhu , Yating Wang , Di Huang , Weicai Ye , Wanli Ouyang , Tong He

We present an approach for aggregating a sparse set of views of an object in order to compute a semi-implicit 3D representation in the form of a volumetric feature grid. Key to our approach is an object-centric canonical 3D coordinate…

计算机视觉与模式识别 · 计算机科学 2020-07-22 Shubham Tulsiani , Or Litany , Charles R. Qi , He Wang , Leonidas J. Guibas

Behavior cloning has shown promise for robot manipulation, but real-world demonstrations are costly to acquire at scale. While simulated data offers a scalable alternative, particularly with advances in automated demonstration generation,…

机器人学 · 计算机科学 2026-01-19 Shuo Cheng , Liqian Ma , Zhenyang Chen , Ajay Mandlekar , Caelan Garrett , Danfei Xu

Equivariant neural networks offer strong inductive biases for learning from molecular and geometric data but often rely on specialized, computationally expensive tensor operations. We present a framework to transfers existing tensor field…

机器学习 · 计算机科学 2025-10-01 Gerrit Gerhartz , Peter Lippmann , Fred A. Hamprecht

Learning for manipulation requires using policies that have access to rich sensory information such as point clouds or RGB images. Point clouds efficiently capture geometric structures, making them essential for manipulation tasks in…

The significance of informative and robust point representations has been widely acknowledged for 3D scene understanding. Despite existing self-supervised pre-training counterparts demonstrating promising performance, the model collapse and…

计算机视觉与模式识别 · 计算机科学 2026-02-12 Lei Yao , Yi Wang , Yi Zhang , Moyun Liu , Lap-Pui Chau

At the core of self-supervised learning for vision is the idea of learning invariant or equivariant representations with respect to a set of data transformations. This approach, however, introduces strong inductive biases, which can render…

机器学习 · 计算机科学 2024-05-29 Sharut Gupta , Chenyu Wang , Yifei Wang , Tommi Jaakkola , Stefanie Jegelka

3D action recognition was shown to benefit from a covariance representation of the input data (joint 3D positions). A kernel machine feed with such feature is an effective paradigm for 3D action recognition, yielding state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2017-10-05 Jacopo Cavazza , Pietro Morerio , Vittorio Murino

Deep neural network models have achieved remarkable progress in 3D scene understanding while trained in the closed-set setting and with full labels. However, the major bottleneck is that these models do not have the capacity to recognize…

计算机视觉与模式识别 · 计算机科学 2025-02-20 Kangcheng Liu , Yong-Jin Liu , Baoquan Chen

Foundation models have significantly enhanced 2D task performance, and recent works like Bridge3D have successfully applied these models to improve 3D scene understanding through knowledge distillation, marking considerable advancements.…

计算机视觉与模式识别 · 计算机科学 2024-10-18 Zhimin Chen , Liang Yang , Yingwei Li , Longlong Jing , Bing Li

Artificial intelligence for scientific discovery has recently generated significant interest within the machine learning and scientific communities, particularly in the domains of chemistry, biology, and material discovery. For these…

In agricultural automation, inherent occlusion presents a major challenge for robotic harvesting. We propose a novel imitation learning-based viewpoint planning approach to actively adjust camera viewpoint and capture unobstructed images of…

机器人学 · 计算机科学 2025-03-14 Lun Li , Hamidreza Kasaei

Nearly all state of the art vision models are sensitive to image rotations. Existing methods often compensate for missing inductive biases by using augmented training data to learn pseudo-invariances. Alongside the resource demanding data…

计算机视觉与模式识别 · 计算机科学 2023-02-08 Johann Schmidt , Sebastian Stober

Behavior cloning methods for robot learning suffer from poor generalization due to limited data support beyond expert demonstrations. Recent approaches leveraging video prediction models have shown promising results by learning rich…

机器人学 · 计算机科学 2025-11-03 Dohyeok Lee , Jung Min Lee , Munkyung Kim , Seokhun Ju , Jin Woo Koo , Kyungjae Lee , Dohyeong Kim , TaeHyun Cho , Jungwoo Lee

Automating garment manipulation is challenging due to extremely high variability in object configurations. To reduce this intrinsic variation, we introduce the task of "canonicalized-alignment" that simplifies downstream applications by…

机器人学 · 计算机科学 2022-10-19 Alper Canberk , Cheng Chi , Huy Ha , Benjamin Burchfiel , Eric Cousineau , Siyuan Feng , Shuran Song

A truly generalizable approach to rigid segmentation and motion estimation is fundamental to 3D understanding of articulated objects and moving scenes. In view of the closely intertwined relationship between segmentation and motion…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Jia-Xing Zhong , Ta-Ying Cheng , Yuhang He , Kai Lu , Kaichen Zhou , Andrew Markham , Niki Trigoni