English
Related papers

Related papers: DinoComplete: 3D Shape Completion with Distilled S…

200 papers

There is substantial interest in developing artificial intelligence systems to support radiologists across tasks ranging from segmentation to report generation. Existing computed tomography (CT) foundation models have largely focused on…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Rubén Moreno-Aguado , Alba Magallón , Victor Moreno , Yingying Fang , Guang Yang

Existing 3D open-vocabulary scene understanding methods mostly emphasize distilling language features from 2D foundation models into 3D feature fields, but largely overlook the synergy among scene appearance, semantics, and geometry. As a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Guile Wu , David Huang , Bingbing Liu , Dongfeng Bai

Recent advancements in sequence modeling have led to the development of the Mamba architecture, noted for its selective state space approach, offering a promising avenue for efficient long sequence handling. However, its application in 3D…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Shentong Mo

Open-Vocabulary Segmentation (OVS) aims to segment image regions beyond predefined category sets by leveraging semantic descriptions. While CLIP based approaches excel in semantic generalization, they frequently lack the fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Haoxi Zeng , Qiankun Liu , Yi Bin , Haiyue Zhang , Yujuan Ding , Guoqing Wang , Deqiang Ouyang , Heng Tao Shen

In this study, we present a multimodal framework for predicting neuro-facial disorders by capturing both vocal and facial cues. We hypothesize that explicitly disentangling shared and modality-specific representations within multimodal…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-13 Mohd Mujtaba Akhtar , Girish , Muskaan Singh

In recent years, Denoising Diffusion Models have demonstrated remarkable success in generating semantically valuable pixel-wise representations for image generative modeling. In this study, we propose a novel end-to-end framework, called…

Image and Video Processing · Electrical Eng. & Systems 2023-03-21 Zhaohu Xing , Liang Wan , Huazhu Fu , Guang Yang , Lei Zhu

Multi-modality image fusion aims to synthesize a single, comprehensive image from multiple source inputs. Traditional approaches, such as CNNs and GANs, offer efficiency but struggle to handle low-quality or complex inputs. Recent advances…

Computer Vision and Pattern Recognition · Computer Science 2025-05-27 Ran Zhang , Xuanhua He , Ke Cao , Liu Liu , Li Zhang , Man Zhou , Jie Zhang

Visual complexity prediction is a fundamental problem in computer vision with applications in image compression, retrieval, and classification. Understanding what makes humans perceive an image as complex is also a long-standing question in…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Jonathan Skaza , Parsa Madinei , Ziqi Wen , Miguel Eckstein

In this paper, we investigate the use of diffusion models which are pre-trained on large-scale image-caption pairs for open-vocabulary 3D semantic understanding. We propose a novel method, namely Diff2Scene, which leverages frozen…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Xiaoyu Zhu , Hao Zhou , Pengfei Xing , Long Zhao , Hao Xu , Junwei Liang , Alexander Hauptmann , Ting Liu , Andrew Gallagher

Semantic scene completion (SSC) aims to complete a partial 3D scene and predict its semantics simultaneously. Most existing works adopt the voxel representations, thus suffering from the growth of memory and computation cost as the voxel…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Jinfeng Xu , Xianzhi Li , Yuan Tang , Qiao Yu , Yixue Hao , Long Hu , Min Chen

In the intersection of computer vision and robotic perception, 4D reconstruction of dynamic scenes serve as the critical bridge connecting low-level geometric sensing with high-level semantic understanding. We present DINO\_4D, introducing…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Yiru Yang , Zhuojie Wu , Quentin Marguet , Nishant Kumar Singh , Max Schulthess

Point cloud completion aims to generate a complete and high-fidelity point cloud from an initially incomplete and low-quality input. A prevalent strategy involves leveraging Transformer-based models to encode global features and facilitate…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Yixuan Li , Weidong Yang , Ben Fei

Semantic scene completion is the task of producing a complete 3D voxel representation of volumetric occupancy with semantic labels for a scene from a single-view observation. We built upon the recent work of Song et al. (CVPR 2017), who…

Computer Vision and Pattern Recognition · Computer Science 2018-02-14 Andre Bernardes Soares Guedes , Teofilo Emidio de Campos , Adrian Hilton

Depth completion involves predicting dense depth maps from sparse LiDAR inputs. However, sparse depth annotations from sensors limit the availability of dense supervision, which is necessary for learning detailed geometric features. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Yingping Liang , Yutao Hu , Wenqi Shao , Ying Fu

Object-centric understanding is fundamental to human vision and required for complex reasoning. Traditional methods define slot-based bottlenecks to learn object properties explicitly, while recent self-supervised vision models like DINO…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Stefan Sylvius Wagner , Stefan Harmeling

Medical image segmentation is critical for diagnosing and treating spinal disorders. However, the presence of high noise, ambiguity, and uncertainty makes this task highly challenging. Factors such as unclear anatomical boundaries,…

Image and Video Processing · Electrical Eng. & Systems 2023-09-13 Zhiqing Zhang , Guojia Fan , Tianyong Liu , Nan Li , Yuyang Liu , Ziyu Liu , Canwei Dong , Shoujun Zhou

Dense 3D shape correspondence remains a central challenge in computer vision and graphics as many deep learning approaches still rely on intermediate geometric features or handcrafted descriptors, limiting their effectiveness under…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Maolin Gao , Shao Jie Hu-Chen , Congyue Deng , Riccardo Marin , Leonidas Guibas , Daniel Cremers

We present a framework that adapts 2D diffusion models for 3D shape completion from incomplete point clouds. While text-to-image diffusion models have achieved remarkable success with abundant 2D data, 3D diffusion models lag due to the…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Yao He , Youngjoong Kwon , Tiange Xiang , Wenxiao Cai , Ehsan Adeli

Training deep models for semantic scene completion (SSC) is challenging due to the sparse and incomplete input, a large quantity of objects of diverse scales as well as the inherent label noise for moving objects. To address the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-14 Zhaoyang Xia , Youquan Liu , Xin Li , Xinge Zhu , Yuexin Ma , Yikang Li , Yuenan Hou , Yu Qiao

Semantic scene completion (SSC) aims to infer both the 3D geometry and semantics of a scene from single images. In contrast to prior work on SSC that heavily relies on expensive ground-truth annotations, we approach SSC in an unsupervised…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Aleksandar Jevtić , Christoph Reich , Felix Wimbauer , Oliver Hahn , Christian Rupprecht , Stefan Roth , Daniel Cremers