中文
相关论文

相关论文: DinoComplete: 3D Shape Completion with Distilled S…

200 篇论文

Medical image registration is a critical component of clinical imaging workflows, enabling accurate longitudinal assessment, multi-modal data fusion, and image-guided interventions. Intensity-based approaches often struggle with…

计算机视觉与模式识别 · 计算机科学 2026-03-17 Eytan Kats , Mattias P. Heinrich

Foundation features from self-supervised vision models and text-to-image diffusion models have proven effective for semantic correspondence estimation. However, because these features are learned primarily from 2D image objectives, they…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Artur Jesslen , Olaf Dünkel , Adam Kortylewski

This paper proposes a cross-modal distillation framework, PartDistill, which transfers 2D knowledge from vision-language models (VLMs) to facilitate 3D shape part segmentation. PartDistill addresses three major challenges in this task: the…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Ardian Umam , Cheng-Kun Yang , Min-Hung Chen , Jen-Hui Chuang , Yen-Yu Lin

We propose a probabilistic shape completion method extended to the continuous geometry of large-scale 3D scenes. Real-world scans of 3D scenes suffer from a considerable amount of missing data cluttered with unsegmented objects. The problem…

计算机视觉与模式识别 · 计算机科学 2022-04-05 Dongsu Zhang , Changwoon Choi , Inbum Park , Young Min Kim

The task of shape abstraction with semantic part consistency is challenging due to the complex geometries of natural objects. Recent methods learn to represent an object shape using a set of simple primitives to fit the target.…

计算机视觉与模式识别 · 计算机科学 2023-09-06 Di Liu , Long Zhao , Qilong Zhangli , Yunhe Gao , Ting Liu , Dimitris N. Metaxas

Accurately determining salient regions of an image is challenging when labeled data is scarce. DINO-based self-supervised approaches have recently leveraged meaningful image semantics captured by patch-wise features for locating foreground…

计算机视觉与模式识别 · 计算机科学 2023-09-21 Sriram Ravindran , Debraj Basu

Vision foundation models have shown great promise for open-set 3D object retrieval (3DOR) through efficient adaptation to multi-view images. Leveraging semantically aligned latent space, previous work typically adapts the CLIP encoder to…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Xinwei He , Yansong Zheng , Qianru Han , Zhichuan Wang , Yuxuan Cai , Yang Zhou , Jingbo Xia , Yulong Wang , Jinhai Xiang , Xiang Bai

Shape reconstruction from imaging volumes is a recurring need in medical image analysis. Common workflows start with a segmentation step, followed by careful post-processing and,finally, ad hoc meshing algorithms. As this sequence can be…

计算机视觉与模式识别 · 计算机科学 2025-07-11 Antonio Pepe , Richard Schussnig , Jianning Li , Christina Gsaxner , Dieter Schmalstieg , Jan Egger

Recovering clean and accurate geometry from images is essential for robotics and augmented reality. However, existing geometry foundation models still suffer severely from flying pixels and the loss of fine details. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Gangwei Xu , Haotong Lin , Hongcheng Luo , Haiyang Sun , Bing Wang , Guang Chen , Sida Peng , Hangjun Ye , Xin Yang

We present a novel 3D shape completion framework that unifies multimodal conditioning, leveraging both 2D images and 3D partial scans through a latent diffusion model. Shapes are represented as Truncated Signed Distance Functions (TSDFs)…

计算机视觉与模式识别 · 计算机科学 2025-09-03 Simon Schaefer , Juan D. Galvis , Xingxing Zuo , Stefan Leutengger

Point clouds collected from real-world environments are often incomplete due to factors such as limited sensor resolution, single viewpoints, occlusions, and noise. These challenges make point cloud completion essential for various…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Yifan Yang , Yuxiang Yan , Boda Liu , Jian Pu

Existing diffusion-based 3D shape completion methods typically use a conditional paradigm, injecting incomplete shape information into the denoising network via deep feature interactions (e.g., concatenation, cross-attention) to guide…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Dequan Kong , Honghua Chen , Zhe Zhu , Mingqiang Wei

Vision foundation models (VFMs) trained on large-scale image datasets provide high-quality features that have significantly advanced 2D visual recognition. However, their potential in 3D scene segmentation remains largely untapped, despite…

计算机视觉与模式识别 · 计算机科学 2026-01-23 Karim Knaebel , Kadir Yilmaz , Daan de Geus , Alexander Hermans , David Adrian , Timm Linder , Bastian Leibe

Vision foundation models (VFMs) such as DINO have led to a paradigm shift in 2D camera-based perception towards extracting generalized features to support many downstream tasks. Recent works introduce self-supervised cross-modal knowledge…

计算机视觉与模式识别 · 计算机科学 2025-11-24 Hariprasath Govindarajan , Maciej K. Wozniak , Marvin Klingner , Camille Maurice , B Ravi Kiran , Senthil Yogamani

Recent works on generalizable NeRFs have shown promising results on novel view synthesis from single or few images. However, such models have rarely been applied on other downstream tasks beyond synthesis such as semantic understanding and…

计算机视觉与模式识别 · 计算机科学 2024-04-10 Jianglong Ye , Naiyan Wang , Xiaolong Wang

Semantic distillation in radiance fields has spurred significant advances in open-vocabulary robot policies, e.g., in manipulation and navigation, founded on pretrained semantics from large vision models. While prior work has demonstrated…

计算机视觉与模式识别 · 计算机科学 2025-10-06 Zhiting Mei , Ola Shorinwa , Anirudha Majumdar

Reconstructing 3D human body shapes from 3D partial textured scans remains a fundamental task for many computer vision and graphics applications -- e.g., body animation, and virtual dressing. We propose a new neural network architecture for…

计算机视觉与模式识别 · 计算机科学 2022-08-23 Ahmet Serdar Karadeniz , Sk Aziz Ali , Anis Kacem , Elona Dupont , Djamila Aouada

This paper introduces VisHall3D, a novel two-stage framework for monocular semantic scene completion that aims to address the issues of feature entanglement and geometric inconsistency prevalent in existing methods. VisHall3D decomposes the…

计算机视觉与模式识别 · 计算机科学 2025-07-28 Haoang Lu , Yuanqi Su , Xiaoning Zhang , Longjun Gao , Yu Xue , Le Wang

Recently, horizontal representation-based panoramic semantic segmentation approaches outperform projection-based solutions, because the distortions can be effectively removed by compressing the spherical data in the vertical direction.…

计算机视觉与模式识别 · 计算机科学 2022-07-07 Zishuo Zheng , Chunyu Lin , Lang Nie , Kang Liao , Zhijie Shen , Yao Zhao

Recently, with the advent of deep convolutional neural networks (DCNN), the improvements in visual saliency prediction research are impressive. One possible direction to approach the next improvement is to fully characterize the multi-scale…

计算机视觉与模式识别 · 计算机科学 2019-05-10 Sheng Yang , Guosheng Lin , Qiuping Jiang , Weisi Lin