English
Related papers

Related papers: NeuroSeg Meets DINOv3: Transferring 2D Self-Superv…

200 papers

Cytoarchitectonic mapping provides anatomically grounded parcellations of brain structure and forms a foundation for integrative, multi-modal neuroscience analyses. These parcellations are defined based on the shape, density, and spatial…

Image and Video Processing · Electrical Eng. & Systems 2026-01-16 Shiqi Zhang , Fang Xu , Pengcheng Zhou

The DINO family of self-supervised vision models has shown remarkable transferability, yet effectively adapting their representations for segmentation remains challenging. Existing approaches often rely on heavy decoders with multi-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Sicheng Yang , Hongqiu Wang , Zhaohu Xing , Sixiang Chen , Lei Zhu

Due to the computational complexity of 3D medical image segmentation, training with downsampled images is a common remedy for out-of-memory errors in deep learning. Nevertheless, as standard spatial convolution is sensitive to variations in…

Image and Video Processing · Electrical Eng. & Systems 2023-10-09 Ken C. L. Wong , Hongzhi Wang , Tanveer Syeda-Mahmood

Neural 3D reconstruction from multi-view images has recently attracted increasing attention from the community. Existing methods normally learn a neural field for the whole scene, while it is still under-explored how to reconstruct a target…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Xiaobao Wei , Renrui Zhang , Jiarui Wu , Jiaming Liu , Ming Lu , Yandong Guo , Shanghang Zhang

Purpose: Depth estimation in robotic surgery is vital in 3D reconstruction, surgical navigation and augmented reality visualization. Although the foundation model exhibits outstanding performance in many vision tasks, including depth…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Beilei Cui , Mobarakol Islam , Long Bai , Hongliang Ren

The difficulties in both data acquisition and annotation substantially restrict the sample sizes of training datasets for 3D medical imaging applications. As a result, constructing high-performance 3D convolutional neural networks from…

Image and Video Processing · Electrical Eng. & Systems 2022-01-06 Shu Zhang , Zihao Li , Hong-Yu Zhou , Jiechao Ma , Yizhou Yu

We present a volume rendering-based neural surface reconstruction method that takes as few as three disparate RGB images as input. Our key idea is to regularize the reconstruction, which is severely ill-posed and leaving significant gaps…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Aditya Vora , Akshay Gadi Patil , Hao Zhang

Medical image analysis frequently encounters data scarcity challenges. Transfer learning has been effective in addressing this issue while conserving computational resources. The recent advent of foundational models like the DINOv2, which…

Image and Video Processing · Electrical Eng. & Systems 2024-02-14 Yuning Huang , Jingchen Zou , Lanxi Meng , Xin Yue , Qing Zhao , Jianqiang Li , Changwei Song , Gabriel Jimenez , Shaowu Li , Guanghui Fu

Foundation vision models are increasingly adopted in medical image analysis. Due to domain shift, these pretrained models misalign with medical image segmentation needs without being fully fine-tuned or lightly adapted. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zhuonan Liang , Wei Guo , Jie Gan , Yaxuan Song , Runnan Chen , Hang Chang , Weidong Cai

Neuroimaging to neuropathology correlation (NTNC) promises to enable the transfer of microscopic signatures of pathology to in vivo imaging with MRI, ultimately enhancing clinical care. NTNC traditionally requires a volumetric MRI scan,…

Instance segmentation enables the analysis of spatial and temporal properties of cells in microscopy images by identifying the pixels belonging to each cell. However, progress is constrained by the scarcity of high-quality labeled…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Kaden Stillwagon , Alexandra Dunnum VandeLoo , Benjamin Magondu , Craig R. Forest

Neuromorphic cameras, also known as event cameras, are asynchronous brightness-change sensors that can capture extremely fast motion without suffering from motion blur, making them particularly promising for 3D reconstruction in extreme…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Chuanzhi Xu , Langyi Chen , Haodong Chen , Vera Chung , Qiang Qu

Adapting foundation models to medical segmentation typically requires either backbone fine-tuning or high-capacity task-specific decoders, both of which are difficult to fit reliably when annotations are scarce. We show that frozen DINOv3…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Wei Jiang , Feng Liu , Nan Ye , Hongfu Sun

This paper evaluates DINOv3, a recent large-scale self-supervised vision backbone, for visuomotor diffusion policy learning in robotic manipulation. We investigate whether a purely self-supervised encoder can match or surpass conventional…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 ThankGod Egbe , Peng Wang , Zhihao Guo , Zidong Chen

Decoding visual information from electroencephalography (EEG) has recently achieved promising results, primarily focusing on reconstructing two-dimensional (2D) images from brain activity. However, the reconstruction of three-dimensional…

Computer Vision and Pattern Recognition · Computer Science 2026-04-10 Emanuele Balloni , Emanuele Frontoni , Chiara Matti , Marina Paolanti , Roberto Pierdicca , Emiliano Santarnecchi

Open-Vocabulary Semantic Segmentation (OVSS) assigns pixel-level labels from an open set of text-defined categories, demanding reliable generalization to unseen classes at inference. Although modern vision-language models (VLMs) support…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Saikat Dutta , Biplab Banerjee , Hamid Rezatofighi

3D shape completion from partial scans remains challenging for unseen categories and noisy real-world observations, where geometry alone is often insufficient for inferring missing structure. We present DinoComplete, a deterministic and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Furkan Mert Algan , Eckehard Steinbach

Medical image registration is a critical component of clinical imaging workflows, enabling accurate longitudinal assessment, multi-modal data fusion, and image-guided interventions. Intensity-based approaches often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Eytan Kats , Mattias P. Heinrich

Current self-supervised learning methods for 3D medical imaging rely on simple pretext formulations and organ- or modality-specific datasets, limiting their generalizability and scalability. We present 3DINO, a cutting-edge SSL method…

Image and Video Processing · Electrical Eng. & Systems 2025-01-22 Tony Xu , Sepehr Hosseini , Chris Anderson , Anthony Rinaldi , Rahul G. Krishnan , Anne L. Martel , Maged Goubran

Vision foundation models (VFMs) trained on large-scale image datasets provide high-quality features that have significantly advanced 2D visual recognition. However, their potential in 3D scene segmentation remains largely untapped, despite…

Computer Vision and Pattern Recognition · Computer Science 2026-01-23 Karim Knaebel , Kadir Yilmaz , Daan de Geus , Alexander Hermans , David Adrian , Timm Linder , Bastian Leibe