English
Related papers

Related papers: DINOv3 Beats Specialized Detectors: A Simple Found…

200 papers

Vision foundation models like DINOv2 demonstrate remarkable potential in medical imaging despite their origin in natural image domains. However, their design inherently works best for uni-modal image analysis, limiting their effectiveness…

Image and Video Processing · Electrical Eng. & Systems 2025-09-09 Daniel Scholz , Ayhan Can Erdur , Viktoria Ehm , Anke Meyer-Baese , Jan C. Peeken , Daniel Rueckert , Benedikt Wiestler

Deep learning models in satellite onboard enable real-time interpretation of remote sensing images, reducing the need for data transmission to the ground and conserving communication resources. As satellite numbers and observation…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Xinyang Pu , Feng Xu

We present LoD-Loc v3, a novel method for generalized aerial visual localization in dense urban environments. While prior work LoD-Loc v2 achieves localization through semantic building silhouette alignment with low-detail city models, it…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Shuaibang Peng , Juelin Zhu , Xia Li , Kun Yang , Maojun Zhang , Yu Liu , Shen Yan

Existing medical image registration algorithms rely on either dataset specific training or local texture-based features to align images. The former cannot be reliably implemented without large modality-specific training datasets, while the…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Xinrui Song , Xuanang Xu , Pingkun Yan

Vision-language models, such as CLIP, have achieved significant success in aligning visual and textual representations, becoming essential components of many multi-modal large language models (MLLMs) like LLaVA and OpenFlamingo. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Shizhan Gong , Yankai Jiang , Qi Dou , Farzan Farnia

Object detection, a crucial aspect of computer vision, has seen significant advancements in accuracy and robustness. Despite these advancements, practical applications still face notable challenges, primarily the inaccurate detection or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Chun-Lin Ji , Tao Yu , Peng Gao , Fei Wang , Ru-Yue Yuan

This paper introduces a novel mathematical property applicable to diverse images, referred to as FINOLA (First-Order Norm+Linear Autoregressive). FINOLA represents each image in the latent space as a first-order autoregressive process, in…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Yinpeng Chen , Xiyang Dai , Dongdong Chen , Mengchen Liu , Lu Yuan , Zicheng Liu , Youzuo Lin

Though current object detection models based on deep learning have achieved excellent results on many conventional benchmark datasets, their performance will dramatically decline on real-world images taken under extreme conditions. Existing…

Computer Vision and Pattern Recognition · Computer Science 2024-06-19 Yuexiong Ding , Xiaowei Luo

Foundation models such as Wav2Vec2 excel at representation learning in speech tasks, including audio deepfake detection. However, after being fine-tuned on a fixed set of bonafide and spoofed audio clips, they often fail to generalize to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-18 Janne Laakkonen , Ivan Kukanov , Ville Hautamäki

Accurate tissue motion tracking is critical to ensure treatment outcome and safety in 2D-Cine MRI-guided radiotherapy. This is typically achieved by registration of sequential images, but existing methods often face challenges with large…

Image and Video Processing · Electrical Eng. & Systems 2025-08-15 Soorena Salari , Catherine Spino , Laurie-Anne Pharand , Fabienne Lathuiliere , Hassan Rivaz , Silvain Beriault , Yiming Xiao

Deep learning-based denoiser has been the focus of recent development on image denoising. In the past few years, there has been increasing interest in developing self-supervised denoising networks that only require noisy images, without the…

Image and Video Processing · Electrical Eng. & Systems 2024-03-20 Jintong Hu , Bin Xia , Bingchen Li , Wenming Yang

Few-shot learning in image classification aims to learn a classifier to classify images when only few training examples are available for each class. Recent work has achieved promising classification performance, where an image-level…

Computer Vision and Pattern Recognition · Computer Science 2019-04-11 Wenbin Li , Lei Wang , Jinglin Xu , Jing Huo , Yang Gao , Jiebo Luo

Recent advances in diffusion-based generative models have demonstrated significant potential in augmenting scarce datasets for object detection tasks. Nevertheless, most recent models rely on resource-intensive full fine-tuning of…

Computer Vision and Pattern Recognition · Computer Science 2025-09-01 Alvaro Patricio , Atabak Dehban , Rodrigo Ventura

Achieving constant accuracy in object detection is challenging due to the inherent variability of object sizes. One effective approach to this problem involves optimizing input resolution, referred to as a multi-resolution strategy.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Daeun Seo , Hoeseok Yang , Hyungshin Kim

With imaging devices delivering ever-higher resolutions and the emerging diffusion-based forgery methods, current detectors trained only on traditional datasets (with splicing, copy-moving and object removal forgeries) lack exposure to this…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Jinhan Li , Haoyang He , Lei Xie , Jiangning Zhang

Advances in high resolution remote sensing image analysis are currently hampered by the difficulty of gathering enough annotated data for training deep learning methods, giving rise to a variety of small datasets and associated…

Computer Vision and Pattern Recognition · Computer Science 2023-03-08 Dimitri Gominski , Valérie Gouet-Brunet , Liming Chen

The emergence of advanced AI-based tools to generate realistic images poses significant challenges for forensic detection and source attribution, especially as new generative techniques appear rapidly. Traditional methods often fail to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Tai D. Nguyen , Aref Azizpour , Matthew C. Stamm

Color ambient lighting normalization under multi-colored illumination is challenging due to severe chromatic shifts, highlight saturation, and material-dependent reflectance. Existing geometric and low-level priors are insufficient for…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Rong-Lin Jian , Ting-Yao Chen , Yu-Fan Lin , Chia-Ming Lee , Fu-En Yang , Yu-Chiang Frank Wang , Chih-Chung Hsu

Foundation models are becoming increasingly popular due to their strong generalization capabilities resulting from being trained on huge datasets. These generalization capabilities are attractive in areas such as NIR Iris Presentation…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Juan E. Tapia , Lázaro Janier González-Soler , Christoph Busch

We present a real-time visual-inertial dense mapping method capable of performing incremental 3D mesh reconstruction with high quality using only sequential monocular images and inertial measurement unit (IMU) readings. 6-DoF camera poses…

Computer Vision and Pattern Recognition · Computer Science 2023-08-29 Yingye Xin , Xingxing Zuo , Dongyue Lu , Stefan Leutenegger