English
Related papers

Related papers: DinoComplete: 3D Shape Completion with Distilled S…

200 papers

Medical image registration is a critical component of clinical imaging workflows, enabling accurate longitudinal assessment, multi-modal data fusion, and image-guided interventions. Intensity-based approaches often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Eytan Kats , Mattias P. Heinrich

Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation. However, because these features are learned primarily from 2D image objectives, they…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Artur Jesslen , Olaf Dünkel , Adam Kortylewski

This paper proposes a cross-modal distillation framework, PartDistill, which transfers 2D knowledge from vision-language models (VLMs) to facilitate 3D shape part segmentation. PartDistill addresses three major challenges in this task: the…

Computer Vision and Pattern Recognition · Computer Science 2024-04-17 Ardian Umam , Cheng-Kun Yang , Min-Hung Chen , Jen-Hui Chuang , Yen-Yu Lin

We propose a probabilistic shape completion method extended to the continuous geometry of large-scale 3D scenes. Real-world scans of 3D scenes suffer from a considerable amount of missing data cluttered with unsegmented objects. The problem…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Dongsu Zhang , Changwoon Choi , Inbum Park , Young Min Kim

The task of shape abstraction with semantic part consistency is challenging due to the complex geometries of natural objects. Recent methods learn to represent an object shape using a set of simple primitives to fit the target.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-06 Di Liu , Long Zhao , Qilong Zhangli , Yunhe Gao , Ting Liu , Dimitris N. Metaxas

Accurately determining salient regions of an image is challenging when labeled data is scarce. DINO-based self-supervised approaches have recently leveraged meaningful image semantics captured by patch-wise features for locating foreground…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Sriram Ravindran , Debraj Basu

Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Xinwei He , Yansong Zheng , Qianru Han , Zhichuan Wang , Yuxuan Cai , Yang Zhou , Jingbo Xia , Yulong Wang , Jinhai Xiang , Xiang Bai

Shape reconstruction from imaging volumes is a recurring need in medical image analysis. Common workflows start with a segmentation step, followed by careful post-processing and,finally, ad hoc meshing algorithms. As this sequence can be…

Computer Vision and Pattern Recognition · Computer Science 2025-07-11 Antonio Pepe , Richard Schussnig , Jianning Li , Christina Gsaxner , Dieter Schmalstieg , Jan Egger

Recovering clean and accurate geometry from images is essential for robotics and augmented reality. However, existing geometry foundation models still suffer severely from flying pixels and the loss of fine details. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2026-01-09 Gangwei Xu , Haotong Lin , Hongcheng Luo , Haiyang Sun , Bing Wang , Guang Chen , Sida Peng , Hangjun Ye , Xin Yang

We present a novel 3D shape completion framework that unifies multimodal conditioning, leveraging both 2D images and 3D partial scans through a latent diffusion model. Shapes are represented as Truncated Signed Distance Functions (TSDFs)…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Simon Schaefer , Juan D. Galvis , Xingxing Zuo , Stefan Leutengger

Point clouds collected from real-world environments are often incomplete due to factors such as limited sensor resolution, single viewpoints, occlusions, and noise. These challenges make point cloud completion essential for various…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Yifan Yang , Yuxiang Yan , Boda Liu , Jian Pu

Existing diffusion-based 3D shape completion methods typically use a conditional paradigm, injecting incomplete shape information into the denoising network via deep feature interactions (e.g., concatenation, cross-attention) to guide…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Dequan Kong , Honghua Chen , Zhe Zhu , Mingqiang Wei

Vision foundation models (VFMs) trained on large-scale image datasets provide high-quality features that have significantly advanced 2D visual recognition. However, their potential in 3D scene segmentation remains largely untapped, despite…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Karim Knaebel , Kadir Yilmaz , Daan de Geus , Alexander Hermans , David Adrian , Timm Linder , Bastian Leibe

Vision foundation models (VFMs) such as DINO have led to a paradigm shift in 2D camera-based perception towards extracting generalized features to support many downstream tasks. Recent works introduce self-supervised cross-modal knowledge…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Hariprasath Govindarajan , Maciej K. Wozniak , Marvin Klingner , Camille Maurice , B Ravi Kiran , Senthil Yogamani

Recent works on generalizable NeRFs have shown promising results on novel view synthesis from single or few images. However, such models have rarely been applied on other downstream tasks beyond synthesis such as semantic understanding and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Jianglong Ye , Naiyan Wang , Xiaolong Wang

Semantic distillation in radiance fields has spurred significant advances in open-vocabulary robot policies, e.g., in manipulation and navigation, founded on pretrained semantics from large vision models. While prior work has demonstrated…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Zhiting Mei , Ola Shorinwa , Anirudha Majumdar

Reconstructing 3D human body shapes from 3D partial textured scans remains a fundamental task for many computer vision and graphics applications -- e.g., body animation, and virtual dressing. We propose a new neural network architecture for…

Computer Vision and Pattern Recognition · Computer Science 2022-08-23 Ahmet Serdar Karadeniz , Sk Aziz Ali , Anis Kacem , Elona Dupont , Djamila Aouada

This paper introduces VisHall3D, a novel two-stage framework for monocular semantic scene completion that aims to address the issues of feature entanglement and geometric inconsistency prevalent in existing methods. VisHall3D decomposes the…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Haoang Lu , Yuanqi Su , Xiaoning Zhang , Longjun Gao , Yu Xue , Le Wang

Recently, horizontal representation-based panoramic semantic segmentation approaches outperform projection-based solutions, because the distortions can be effectively removed by compressing the spherical data in the vertical direction.…

Computer Vision and Pattern Recognition · Computer Science 2022-07-07 Zishuo Zheng , Chunyu Lin , Lang Nie , Kang Liao , Zhijie Shen , Yao Zhao

Recently, with the advent of deep convolutional neural networks (DCNN), the improvements in visual saliency prediction research are impressive. One possible direction to approach the next improvement is to fully characterize the multi-scale…

Computer Vision and Pattern Recognition · Computer Science 2019-05-10 Sheng Yang , Guosheng Lin , Qiuping Jiang , Weisi Lin