中文
相关论文

相关论文: DINO-BOLDNet: A DINOv3-Guided Multi-Slice Attentio…

200 篇论文

Transformers trained with self-supervised learning using self-distillation loss (DINO) have been shown to produce attention maps that highlight salient foreground objects. In this paper, we demonstrate a graph-based approach that uses the…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Yangtao Wang , Xi Shen , Shell Hu , Yuan Yuan , James Crowley , Dominique Vaufreydaz

State-of-the-art vessel segmentation methods typically require large-scale annotated datasets and suffer from severe performance degradation under domain shifts. In clinical practice, however, acquiring extensive annotations for every new…

图像与视频处理 · 电气工程与系统科学 2026-03-02 Kirato Yoshihara , Yohei Sugawara , Yuta Tokuoka , Lihang Hong

This paper evaluates DINOv3, a recent large-scale self-supervised vision backbone, for visuomotor diffusion policy learning in robotic manipulation. We investigate whether a purely self-supervised encoder can match or surpass conventional…

计算机视觉与模式识别 · 计算机科学 2025-09-23 ThankGod Egbe , Peng Wang , Zhihao Guo , Zidong Chen

2D visual foundation models, such as DINOv3, a self-supervised model trained on large-scale natural images, have demonstrated strong zero-shot generalization, capturing both rich global context and fine-grained structural cues. However, an…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Yik San Cheng , Runkai Zhao , Weidong Cai

Scientific machine learning has enabled the extraction of physical insights and data-driven modeling of high-dimensional spatiotemporal data, yet achieving physically interpretable latent representations and computationally efficient…

机器学习 · 计算机科学 2026-05-04 Siva Viknesh , Amirhossein Arzani

Recent self-supervised Vision Transformers (ViTs), such as DINOv3, provide rich feature representations for dense vision tasks. This study investigates the intrinsic few-shot semantic segmentation (FSS) capabilities of frozen DINOv3…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Hussni Mohd Zakir , Eric Tatt Wei Ho

Although purely transformer-based architectures showed promising performance in many computer vision tasks, many hybrid models consisting of CNN and transformer blocks are introduced to fit more specialized tasks. Nevertheless, despite the…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Yousef Yeganeh , Azade Farshad , Peter Weinberger , Seyed-Ahmad Ahmadi , Ehsan Adeli , Nassir Navab

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this…

Foundation models in artificial intelligence (AI) are transforming medical imaging by enabling general-purpose feature learning from large-scale, unlabeled datasets. In this work, we introduce BrainFound, a self-supervised foundation model…

计算机视觉与模式识别 · 计算机科学 2026-01-01 Moona Mazher , Geoff J. M. Parker , Daniel C. Alexander

The scarcity and high cost of expert annotations in dental imaging present a significant challenge for the development of AI in dentistry. DINOv3, a state-of-the-art, self-supervised vision foundation model pre-trained on 1.7 billion…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Kun Tang , Xinquan Yang , Mianjie Zheng , Xuefen Liu , Xuguang Li , Xiaoqi Guo , Ruihan Chen , Linlin Shen , He Meng

This paper presents DINO-RotateMatch, a deep-learning framework designed to address the chal lenges of image matching in large-scale 3D reconstruction from unstructured Internet images. The method integrates a dataset-adaptive image pairing…

计算机视觉与模式识别 · 计算机科学 2025-12-04 Kaichen Zhang , Tianxiang Sheng , Xuanming Shi

Medical imaging is limited by acquisition time and scanning equipment. CT and MR volumes, reconstructed with thicker slices, are anisotropic with high in-plane resolution and low through-plane resolution. We reveal an intriguing phenomenon…

图像与视频处理 · 电气工程与系统科学 2024-05-07 Haofei Song , Xintian Mao , Jing Yu , Qingli Li , Yan Wang

Deep learning models have emerged as the cornerstone of medical image segmentation, but their efficacy hinges on the availability of extensive manually labeled datasets and their adaptability to unforeseen categories remains a challenge.…

计算机视觉与模式识别 · 计算机科学 2024-03-07 Lev Ayzenberg , Raja Giryes , Hayit Greenspan

Accurate segmentation of medical images is crucial for diagnostic purposes, including cell segmentation, tumor identification, and organ localization. Traditional convolutional neural network (CNN)-based approaches struggled to achieve…

图像与视频处理 · 电气工程与系统科学 2024-06-26 Daniya Najiha Abdul Kareem , Mustansar Fiaz , Noa Novershtern , Hisham Cholakkal

In remote sensing imagery, multi class change detection (MCD) is crucial for fine grained monitoring, yet it has long been constrained by complex scene variations and the scarcity of detailed annotations. To address this, we propose the…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Kai Zheng , Hang-Cheng Dong , Shoulei Liu , Zhenkai Wu , Fupeng Wei , Lei Ding , Wei Zhang

With the rapid advancement of deep generative models, realistic fake images have become increasingly accessible, yet existing localization methods rely on complex designs and still struggle to generalize across manipulation types and…

计算机视觉与模式识别 · 计算机科学 2026-04-20 Jieming Yu , Qiuxiao Feng , Zhuohan Wang , Xiaochen Ma

Image denoising aims to restore a clean image from an observed noisy image. The model-based image denoising approaches can achieve good generalization ability over different noise levels and are with high interpretability. Learning-based…

图像与视频处理 · 电气工程与系统科学 2022-07-13 Jun-Jie Huang , Pier Luigi Dragotti

Vision foundation models pretrained on web-scale data have recently shown strong transfer capabilities on many downstream tasks, but their effectiveness for industrial visual inspection remains unclear. Industrial data differ substantially…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Mehdi Gharbage , Céline Teulière , Pierre Bouges , Thierry Chateau

Extracting narrow roads from high-resolution remote sensing imagery remains a significant challenge due to their limited width, fragmented topology, and frequent occlusions. To address these issues, we propose D3FNet, a Dilated Dual-Stream…

计算机视觉与模式识别 · 计算机科学 2025-08-22 Chang Liu , Yang Xu , Tamas Sziranyi

Due to the computational complexity of 3D medical image segmentation, training with downsampled images is a common remedy for out-of-memory errors in deep learning. Nevertheless, as standard spatial convolution is sensitive to variations in…

图像与视频处理 · 电气工程与系统科学 2023-10-09 Ken C. L. Wong , Hongzhi Wang , Tanveer Syeda-Mahmood