中文
相关论文

相关论文: DINOv3 Beats Specialized Detectors: A Simple Found…

200 篇论文

As generative models become increasingly diverse and powerful, cross-generator detection has emerged as a new challenge. Existing detection methods often memorize artifacts of specific generative models rather than learning transferable…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Zhenglin Huang , Jason Li , Haiquan Wen , Tianxiao Li , Xi Yang , Lu Qi , Bei Peng , Xiaowei Huang , Ming-Hsuan Yang , Guangliang Cheng

The scarcity and high cost of expert annotations in dental imaging present a significant challenge for the development of AI in dentistry. DINOv3, a state-of-the-art, self-supervised vision foundation model pre-trained on 1.7 billion…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Kun Tang , Xinquan Yang , Mianjie Zheng , Xuefen Liu , Xuguang Li , Xiaoqi Guo , Ruihan Chen , Linlin Shen , He Meng

Foundation models have become prominent in computer vision, achieving notable success in various tasks. However, their effectiveness largely depends on pre-training with extensive datasets. Applying foundation models directly to small…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Bowen Zhang , Ying Chen , Long Bai , Yan Zhao , Yuxiang Sun , Yixuan Yuan , Jianhua Zhang , Hongliang Ren

Purpose: Depth estimation in robotic surgery is vital in 3D reconstruction, surgical navigation and augmented reality visualization. Although the foundation model exhibits outstanding performance in many vision tasks, including depth…

计算机视觉与模式识别 · 计算机科学 2024-01-15 Beilei Cui , Mobarakol Islam , Long Bai , Hongliang Ren

The advent of large-scale vision foundation models, pre-trained on diverse natural images, has marked a paradigm shift in computer vision. However, how the frontier vision foundation models' efficacies transfer to specialised domains such…

Recent advances in multimodal foundation models have set new standards in few-shot anomaly detection. This paper explores whether high-quality visual features alone are sufficient to rival existing state-of-the-art vision-language models.…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Simon Damm , Mike Laszkiewicz , Johannes Lederer , Asja Fischer

Face Anti-Spoofing (FAS) remains challenging due to the requirement for robust domain generalization across unseen environments. While recent trends leverage Vision-Language Models (VLMs) for semantic supervision, these multimodal…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Mika Feng , Pierre Gallin-Martel , Koichi Ito , Takafumi Aoki

Generalization under manipulation and dataset shift remains a core challenge in forged media detection for AI-driven edge sensing systems. Frozen vision foundation models with linear probes are strong baselines, but most pipelines use…

图像与视频处理 · 电气工程与系统科学 2026-03-30 Izaldein Al-Zyoud , Abdulmotaleb El Saddik

Utilizing visual place recognition (VPR) technology to ascertain the geographical location of publicly available images is a pressing issue for real-world VPR applications. Although most current VPR methods achieve favorable results under…

计算机视觉与模式识别 · 计算机科学 2023-12-06 Gaoshuang Huang , Yang Zhou , Xiaofei Hu , Chenglong Zhang , Luying Zhao , Wenjian Gan , Mingbo Hou

Object detection in civil engineering applications is constrained by limited annotated data in specialized domains. We introduce DINO-YOLO, a hybrid architecture combining YOLOv12 with DINOv3 self-supervised vision transformers for…

计算机视觉与模式识别 · 计算机科学 2025-11-03 Malaisree P , Youwai S , Kitkobsin T , Janrungautai S , Amorndechaphon D , Rojanavasu P

Recent self-supervised Vision Transformers (ViTs), such as DINOv3, provide rich feature representations for dense vision tasks. This study investigates the intrinsic few-shot semantic segmentation (FSS) capabilities of frozen DINOv3…

计算机视觉与模式识别 · 计算机科学 2026-02-10 Hussni Mohd Zakir , Eric Tatt Wei Ho

Foundation models pre-trained on large-scale natural image datasets offer a powerful paradigm for medical image segmentation. However, effectively transferring their learned representations for precise clinical applications remains a…

计算机视觉与模式识别 · 计算机科学 2026-05-11 Haoyue Li , Yifan Gao , Feng Yuan , Xiaosong Wang , Xin Gao

Low-Light Image Enhancement (LLIE) is a key task in computational photography and imaging. The problem of enhancing images captured during night or in dark environments has been well-studied in the computer vision literature. However,…

计算机视觉与模式识别 · 计算机科学 2026-01-29 Juan C. Benito , Daniel Feijoo , Alvaro Garcia , Marcos V. Conde

The increasing accessibility of image editing tools and generative AI has led to a proliferation of visually convincing forgeries, compromising the authenticity of digital media. In this paper, in addition to leveraging distortions from…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Youqi Wang , Shunquan Tan , Rongxuan Peng , Bin Li , Jiwu Huang

Atypical mitotic figures (AMFs) represent abnormal cell division associated with poor prognosis. Yet their detection remains difficult due to low prevalence, subtle morphology, and inter-observer variability. The MIDOG 2025 challenge…

图像与视频处理 · 电气工程与系统科学 2025-10-15 Guillaume Balezo , Hana Feki , Raphaël Bourgade , Lily Monnier , Matthieu Blons , Alice Blondel , Etienne Decencière , Albert Pla Planas , Thomas Walter

Learning-based monocular visual odometry (VO) poses robustness, generalization, and efficiency challenges in robotics. Recent advances in visual foundation models, such as DINOv2, have improved robustness and generalization in various…

计算机视觉与模式识别 · 计算机科学 2025-07-18 Maulana Bisyir Azhari , David Hyunchul Shim

The proliferation of sophisticated deepfakes poses significant threats to information integrity. While DINOv2 shows promise for detection, existing fine-tuning approaches treat it as generic binary classification, overlooking distinct…

计算机视觉与模式识别 · 计算机科学 2025-11-18 Tianxiang Zhang , Peipeng Yu , Zhihua Xia , Longchen Dai , Xiaoyu Zhou , Hui Gao

With growing concerns over image authenticity and digital safety, the field of AI-generated image (AIGI) detection has progressed rapidly. Yet, most AIGI detectors still struggle under real-world degradations, particularly motion blur,…

计算机视觉与模式识别 · 计算机科学 2025-11-19 Jialiang Shen , Jiyang Zheng , Yunqi Xue , Huajie Chen , Yu Yao , Hui Kang , Ruiqi Liu , Helin Gong , Yang Yang , Dadong Wang , Tongliang Liu

Self-supervised learning holds the promise of eliminating the need for manual data annotation, enabling models to scale effortlessly to massive datasets and larger architectures. By not being tailored to specific tasks or domains, this…

2D visual foundation models, such as DINOv3, a self-supervised model trained on large-scale natural images, have demonstrated strong zero-shot generalization, capturing both rich global context and fine-grained structural cues. However, an…

计算机视觉与模式识别 · 计算机科学 2026-03-25 Yik San Cheng , Runkai Zhao , Weidong Cai
‹ 上一页 1 2 3 10 下一页 ›