中文
相关论文

相关论文: Frozen Vision Transformers for Dense Prediction on…

200 篇论文

Vision Transformers (ViTs) can learn strong image-level representations while their patch representations become less effective for dense prediction during prolonged training. We revisit this dense degradation phenomenon and argue that it…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Linxiang Su

Transformer-based architectures have established a dominant paradigm in global semantic perception; however, they remain fundamentally constrained by the profound spatial heterogeneity inherent in natural images. Specifically, the…

计算机视觉与模式识别 · 计算机科学 2026-04-30 Hui Wang , Hongze Li , Wei Chen , Xiaojin Zhang

Transformers trained with self-supervised learning using self-distillation loss (DINO) have been shown to produce attention maps that highlight salient foreground objects. In this paper, we demonstrate a graph-based approach that uses the…

计算机视觉与模式识别 · 计算机科学 2022-03-25 Yangtao Wang , Xi Shen , Shell Hu , Yuan Yuan , James Crowley , Dominique Vaufreydaz

Directed Energy Deposition (DED) offers significant potential for manufacturing complex and multi-material parts. However, internal defects such as porosity and cracks can compromise mechanical properties and overall performance. This study…

计算机视觉与模式识别 · 计算机科学 2025-02-12 Israt Zarin Era , Fan Zhou , Ahmed Shoyeb Raihan , Imtiaz Ahmed , Alan Abul-Haj , James Craig , Srinjoy Das , Zhichao Liu

This work presents a neural network model capable of recognizing small and tiny objects in thermal images collected by unmanned aerial vehicles. Our model consists of three parts, the backbone, the neck, and the prediction head. The…

计算机视觉与模式识别 · 计算机科学 2024-02-14 Minh Dang Tu , Kieu Trang Le , Manh Duong Phung

Rotated object detection in aerial images has received increasing attention for a wide range of applications. However, it is also a challenging task due to the huge variations of scale, rotation, aspect ratio, and densely arranged targets.…

计算机视觉与模式识别 · 计算机科学 2021-11-29 Feng Zhang , Xueying Wang , Shilin Zhou , Yingqian Wang

This work addresses visual cross-view metric localization for outdoor robotics. Given a ground-level color image and a satellite patch that contains the local surroundings, the task is to identify the location of the ground camera within…

计算机视觉与模式识别 · 计算机科学 2022-08-19 Zimin Xia , Olaf Booij , Marco Manfredi , Julian F. P. Kooij

Image dehazing is a representative low-level vision task that estimates latent haze-free images from hazy images. In recent years, convolutional neural network-based methods have dominated image dehazing. However, vision Transformers, which…

计算机视觉与模式识别 · 计算机科学 2023-04-12 Yuda Song , Zhuqing He , Hui Qian , Xin Du

The Vision Transformer architecture is a deep learning model inspired by the success of the Transformer model in Natural Language Processing. However, the self-attention mechanism, large number of parameters, and the requirement for a…

计算机视觉与模式识别 · 计算机科学 2023-07-25 Yogi Prasetyo , Novanto Yudistira , Agus Wahyu Widodo

When deploying Deep Neural Networks (DNNs), developers often convert models from one deep learning framework to another (e.g., TensorFlow to PyTorch). However, this process is error-prone and can impact target model accuracy. To identify…

计算机视觉与模式识别 · 计算机科学 2024-03-27 Nikolaos Louloudakis , Perry Gibson , José Cano , Ajitha Rajan

This paper introduces a self-supervised learning framework designed for pre-training neural networks tailored to dense prediction tasks using event camera data. Our approach utilizes solely event data for training. Transferring achievements…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Yan Yang , Liyuan Pan , Liu Liu

In this paper, we introduce a novel approach to fine-grained cross-view geo-localization. Our method aligns a warped ground image with a corresponding GPS-tagged satellite image covering the same area using homography estimation. We first…

计算机视觉与模式识别 · 计算机科学 2023-09-01 Xiaolong Wang , Runsen Xu , Zuofan Cui , Zeyu Wan , Yu Zhang

The quick and accurate retrieval of an object height from a single fringe pattern in Fringe Projection Profilometry has been a topic of ongoing research. While a single shot fringe to depth CNN based method can restore height map directly…

计算机视觉与模式识别 · 计算机科学 2023-05-01 Yixiao Wang , Canlin Zhou , Xingyang Qi , Hui Li

This paper introduces self-taught object localization, a novel approach that leverages deep convolutional networks trained for whole-image recognition to localize objects in images without additional human supervision, i.e., without using…

计算机视觉与模式识别 · 计算机科学 2016-02-03 Loris Bazzani , Alessandro Bergamo , Dragomir Anguelov , Lorenzo Torresani

Extracting standardized metallurgical metrics from microscopy images remains challenging due to complex grain morphology and the data demands of supervised segmentation. To bridge foundational computer vision with practical metallurgical…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Abdul Mueez , Shruti Vyas

Deploying high-performance dense prediction models on resource-constrained edge devices remains challenging due to strict limits on computation and memory. In practice, lightweight systems for object detection, instance segmentation, and…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Longfei Liu , Yongjie Hou , Yang Li , Qirui Wang , Youyang Sha , Yongjun Yu , Yinzhi Wang , Peizhe Ru , Xuanlong Yu , Xi Shen

Anomaly detection methods typically require extensive normal samples from the target class for training, limiting their applicability in scenarios that require rapid adaptation, such as cold start. Zero-shot and few-shot anomaly detection…

计算机视觉与模式识别 · 计算机科学 2026-03-03 Zhaopeng Gu , Bingke Zhu , Guibo Zhu , Yingying Chen , Ming Tang , Jinqiao Wang

Real-time, high-fidelity monocular depth estimation from remote sensing imagery is crucial for numerous applications, yet existing methods face a stark trade-off between accuracy and efficiency. Although using Vision Transformer (ViT)…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Ruizhi Wang , Weihan Li , Zunlei Feng , Haofei Zhang , Mingli Song , Jiayu Wang , Jie Song , Li Sun

The need for large annotated image datasets for training Convolutional Neural Networks (CNNs) has been a significant impediment for their adoption in computer vision applications. We show that with transfer learning an effective object…

计算机视觉与模式识别 · 计算机科学 2017-09-19 Param S. Rajpura , Hristo Bojinov , Ravi S. Hegde

Precision agriculture involves the application of advanced technologies to improve agricultural productivity, efficiency, and profitability while minimizing waste and environmental impact. Deep learning approaches enable automated…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Alireza Ghanbari , Gholamhassan Shirdel , Farhad Maleki