English
Related papers

Related papers: DINOv3 Beats Specialized Detectors: A Simple Found…

200 papers

As generative models become increasingly diverse and powerful, cross-generator detection has emerged as a new challenge. Existing detection methods often memorize artifacts of specific generative models rather than learning transferable…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Zhenglin Huang , Jason Li , Haiquan Wen , Tianxiao Li , Xi Yang , Lu Qi , Bei Peng , Xiaowei Huang , Ming-Hsuan Yang , Guangliang Cheng

The scarcity and high cost of expert annotations in dental imaging present a significant challenge for the development of AI in dentistry. DINOv3, a state-of-the-art, self-supervised vision foundation model pre-trained on 1.7 billion…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Kun Tang , Xinquan Yang , Mianjie Zheng , Xuefen Liu , Xuguang Li , Xiaoqi Guo , Ruihan Chen , Linlin Shen , He Meng

Foundation models have become prominent in computer vision, achieving notable success in various tasks. However, their effectiveness largely depends on pre-training with extensive datasets. Applying foundation models directly to small…

Computer Vision and Pattern Recognition · Computer Science 2024-07-02 Bowen Zhang , Ying Chen , Long Bai , Yan Zhao , Yuxiang Sun , Yixuan Yuan , Jianhua Zhang , Hongliang Ren

Purpose: Depth estimation in robotic surgery is vital in 3D reconstruction, surgical navigation and augmented reality visualization. Although the foundation model exhibits outstanding performance in many vision tasks, including depth…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Beilei Cui , Mobarakol Islam , Long Bai , Hongliang Ren

The advent of large-scale vision foundation models, pre-trained on diverse natural images, has marked a paradigm shift in computer vision. However, how the frontier vision foundation models' efficacies transfer to specialised domains such…

Recent advances in multimodal foundation models have set new standards in few-shot anomaly detection. This paper explores whether high-quality visual features alone are sufficient to rival existing state-of-the-art vision-language models.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Simon Damm , Mike Laszkiewicz , Johannes Lederer , Asja Fischer

Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Mika Feng , Pierre Gallin-Martel , Koichi Ito , Takafumi Aoki

Generalization under manipulation and dataset shift remains a core challenge in forged media detection for AI-driven edge sensing systems. Frozen vision foundation models with linear probes are strong baselines, but most pipelines use…

Image and Video Processing · Electrical Eng. & Systems 2026-03-30 Izaldein Al-Zyoud , Abdulmotaleb El Saddik

Utilizing visual place recognition (VPR) technology to ascertain the geographical location of publicly available images is a pressing issue for real-world VPR applications. Although most current VPR methods achieve favorable results under…

Computer Vision and Pattern Recognition · Computer Science 2023-12-06 Gaoshuang Huang , Yang Zhou , Xiaofei Hu , Chenglong Zhang , Luying Zhao , Wenjian Gan , Mingbo Hou

Object detection in civil engineering applications is constrained by limited annotated data in specialized domains. We introduce DINO-YOLO, a hybrid architecture combining YOLOv12 with DINOv3 self-supervised vision transformers for…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Malaisree P , Youwai S , Kitkobsin T , Janrungautai S , Amorndechaphon D , Rojanavasu P

Recent self-supervised Vision Transformers (ViTs), such as DINOv3, provide rich feature representations for dense vision tasks. This study investigates the intrinsic few-shot semantic segmentation (FSS) capabilities of frozen DINOv3…

Computer Vision and Pattern Recognition · Computer Science 2026-02-10 Hussni Mohd Zakir , Eric Tatt Wei Ho

Foundation models pre-trained on large-scale natural image datasets offer a powerful paradigm for medical image segmentation. However, effectively transferring their learned representations for precise clinical applications remains a…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Haoyue Li , Yifan Gao , Feng Yuan , Xiaosong Wang , Xin Gao

Low-Light Image Enhancement (LLIE) is a key task in computational photography and imaging. The problem of enhancing images captured during night or in dark environments has been well-studied in the computer vision literature. However,…

Computer Vision and Pattern Recognition · Computer Science 2026-01-29 Juan C. Benito , Daniel Feijoo , Alvaro Garcia , Marcos V. Conde

The increasing accessibility of image editing tools and generative AI has led to a proliferation of visually convincing forgeries, compromising the authenticity of digital media. In this paper, in addition to leveraging distortions from…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Youqi Wang , Shunquan Tan , Rongxuan Peng , Bin Li , Jiwu Huang

Atypical mitotic figures (AMFs) represent abnormal cell division associated with poor prognosis. Yet their detection remains difficult due to low prevalence, subtle morphology, and inter-observer variability. The MIDOG 2025 challenge…

Image and Video Processing · Electrical Eng. & Systems 2025-10-15 Guillaume Balezo , Hana Feki , Raphaël Bourgade , Lily Monnier , Matthieu Blons , Alice Blondel , Etienne Decencière , Albert Pla Planas , Thomas Walter

Learning-based monocular visual odometry (VO) poses robustness, generalization, and efficiency challenges in robotics. Recent advances in visual foundation models, such as DINOv2, have improved robustness and generalization in various…

Computer Vision and Pattern Recognition · Computer Science 2025-07-18 Maulana Bisyir Azhari , David Hyunchul Shim

The proliferation of sophisticated deepfakes poses significant threats to information integrity. While DINOv2 shows promise for detection, existing fine-tuning approaches treat it as generic binary classification, overlooking distinct…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Tianxiang Zhang , Peipeng Yu , Zhihua Xia , Longchen Dai , Xiaoyu Zhou , Hui Gao

With growing concerns over image authenticity and digital safety, the field of AI-generated image (AIGI) detection has progressed rapidly. Yet, most AIGI detectors still struggle under real-world degradations, particularly motion blur,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-19 Jialiang Shen , Jiyang Zheng , Yunqi Xue , Huajie Chen , Yu Yao , Hui Kang , Ruiqi Liu , Helin Gong , Yang Yang , Dadong Wang , Tongliang Liu

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this…

2D visual foundation models, such as DINOv3, a self-supervised model trained on large-scale natural images, have demonstrated strong zero-shot generalization, capturing both rich global context and fine-grained structural cues. However, an…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Yik San Cheng , Runkai Zhao , Weidong Cai
‹ Prev 1 2 3 10 Next ›