中文
相关论文

相关论文: SE(3)-bi-equivariant Transformers for Point Cloud …

200 篇论文

Model binarization has made significant progress in enabling real-time and energy-efficient computation for convolutional neural networks (CNN), offering a potential solution to the deployment challenges faced by Vision Transformers (ViTs)…

计算机视觉与模式识别 · 计算机科学 2025-03-07 Tian Gao , Zhiyuan Zhang , Yu Zhang , Huajun Liu , Kaijie Yin , Chengzhong Xu , Hui Kong

Regular group convolutional neural networks (G-CNNs) have been shown to increase model performance and improve equivariance to different geometrical symmetries. This work addresses the problem of SE(3), i.e., roto-translation equivariance,…

计算机视觉与模式识别 · 计算机科学 2023-07-21 Thijs P. Kuipers , Erik J. Bekkers

Vision Transformer (ViT) has achieved remarkable performance in computer vision. However, positional encoding in ViT makes it substantially difficult to learn the intrinsic equivariance in data. Initial attempts have been made on designing…

计算机视觉与模式识别 · 计算机科学 2023-07-10 Renjun Xu , Kaifan Yang , Ke Liu , Fengxiang He

Point clouds obtained from capture devices or 3D reconstruction techniques are often noisy and interfere with downstream tasks. The paper aims to recover the underlying surface of noisy point clouds. We design a novel model, NoiseTrans,…

计算机视觉与模式识别 · 计算机科学 2023-04-25 Guangzhe Hou , Guihe Qin , Minghui Sun , Yanhua Liang , Jie Yan , Zhonghan Zhang

Shift equivariance is a fundamental principle that governs how we perceive the world - our recognition of an object remains invariant with respect to shifts. Transformers have gained immense popularity due to their effectiveness in both…

计算机视觉与模式识别 · 计算机科学 2023-06-14 Peijian Ding , Davit Soselia , Thomas Armstrong , Jiahao Su , Furong Huang

To enable versatile robot manipulation, robots must detect task-relevant poses for different purposes from raw scenes. Currently, many perception algorithms are designed for specific purposes, which limits the flexibility of the perception…

机器人学 · 计算机科学 2024-11-18 Kanghyun Kim , Min Jun Kim

Test-Time Training (TTT) has emerged as a promising solution to address distribution shifts in 3D point cloud classification. However, existing methods often rely on computationally expensive backpropagation during adaptation, limiting…

We introduce Steerable Transformers, an extension of the Vision Transformer mechanism that maintains equivariance to the special Euclidean group $\mathrm{SE}(d)$. We propose an equivariant attention mechanism that operates on features…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Soumyabrata Kundu , Risi Kondor

This paper proposes a set of rules to revise various neural networks for 3D point cloud processing to rotation-equivariant quaternion neural networks (REQNNs). We find that when a neural network uses quaternion features under certain…

计算机视觉与模式识别 · 计算机科学 2020-10-13 Wen Shen , Binbin Zhang , Shikun Huang , Zhihua Wei , Quanshi Zhang

Recently, 3D understanding research sheds light on extracting features from point cloud directly, which requires effective shape pattern description of point clouds. Inspired by the outstanding 2D shape descriptor SIFT, we design a module…

计算机视觉与模式识别 · 计算机科学 2018-11-27 Mingyang Jiang , Yiran Wu , Tianqi Zhao , Zelin Zhao , Cewu Lu

Unstructured point clouds with varying sizes are increasingly acquired in a variety of environments through laser triangulation or Light Detection and Ranging (LiDAR). Predicting a scalar response based on unstructured point clouds is a…

计算机视觉与模式识别 · 计算机科学 2022-03-01 Michael Biehler , Hao Yan , Jianjun Shi

Existing LiDAR-Camera fusion methods have achieved strong results in 3D object detection. To address the sparsity of point clouds, previous approaches typically construct spatial pseudo point clouds via depth completion as auxiliary input…

计算机视觉与模式识别 · 计算机科学 2025-08-28 Jijun Wang , Yan Wu , Yujian Mo , Junqiao Zhao , Jun Yan , Yinghao Hu

Rigid motion tracking is paramount in many medical imaging applications where movements need to be detected, corrected, or accounted for. Modern strategies rely on convolutional neural networks (CNN) and pose this problem as rigid…

图像与视频处理 · 电气工程与系统科学 2024-06-13 Benjamin Billot , Neel Dey , Daniel Moyer , Malte Hoffmann , Esra Abaci Turk , Borjan Gagoski , Ellen Grant , Polina Golland

We present a surprisingly simple and efficient method for self-supervision of 3D backbone on automotive Lidar point clouds. We design a contrastive loss between features of Lidar scans captured in the same scene. Several such approaches…

计算机视觉与模式识别 · 计算机科学 2023-10-27 Corentin Sautier , Gilles Puy , Alexandre Boulch , Renaud Marlet , Vincent Lepetit

Cross-attention transformers and other multimodal vision-language models excel at grounding and generation; however, their extensive, full-precision backbones make it challenging to deploy them on edge devices. Memory-augmented…

计算与语言 · 计算机科学 2025-10-14 Euhid Aman , Esteban Carlin , Hsing-Kuo Pao , Giovanni Beltrame , Ghaluh Indah Permata Sari , Yie-Tarng Chen

Learning to assemble geometric shapes into a larger target structure is a pivotal task in various practical applications. In this work, we tackle this problem by establishing local correspondences between point clouds of part shapes in both…

计算机视觉与模式识别 · 计算机科学 2024-07-16 Nahyuk Lee , Juhong Min , Junha Lee , Seungwook Kim , Kanghee Lee , Jaesik Park , Minsu Cho

We can use a method called registration to integrate some point clouds that represent the shape of the real world. In this paper, we propose highly accurate and stable registration method. Our method detects keypoints from point clouds and…

计算机视觉与模式识别 · 计算机科学 2020-11-11 Masaki Yoshii , Ikuko Shimizu

Many datasets in scientific and engineering applications are comprised of objects which have specific geometric structure. A common example is data which inhabits a representation of the group SO$(3)$ of 3D rotations: scalars, vectors,…

机器学习 · 计算机科学 2023-03-21 Chase Shimmin , Zhelun Li , Ema Smith

Shape-Text matching is an important task of high-level shape understanding. Current methods mainly represent a 3D shape as multiple 2D rendered views, which obviously can not be understood well due to the structural ambiguity caused by…

计算机视觉与模式识别 · 计算机科学 2023-04-04 Chuan Tang , Xi Yang , Bojian Wu , Zhizhong Han , Yi Chang

Identifying changes in a pair of 3D aerial LiDAR point clouds, obtained during two distinct time periods over the same geographic region presents a significant challenge due to the disparities in spatial coverage and the presence of noise…

计算机视觉与模式识别 · 计算机科学 2023-08-31 Peter Naylor , Diego Di Carlo , Arianna Traviglia , Makoto Yamada , Marco Fiorucci