中文
相关论文

相关论文: Zero-Ablation Overstates Register Content Dependen…

200 篇论文

Real-world visual systems face time-varying perturbations, including weather, sensor noise, compression artifacts, and background distractions. Existing image restoration methods are typically designed for fixed corruption types and…

机器人学 · 计算机科学 2026-05-11 Zhengru Fang , Yu Guo , Fei Liu , Yuang Zhang , Yihang Tao , Senkang Hu , Wenbo Ding , Yuguang Fang

Training-free camera control for pretrained flow-matching video generators is a partial-observation inverse problem: a depth-warped guidance video supplies noisy evidence on a subset of latent sites, which the sampler must reconcile with…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Yuzhu Wang , Xi Ye , Duo Su , Yangyang Xu , Jun Zhu

Denoising, the process of reducing random fluctuations in a signal to emphasize essential patterns, has been a fundamental problem of interest since the dawn of modern scientific inquiry. Recent denoising techniques, particularly in…

机器学习 · 计算机科学 2024-12-04 Peyman Milanfar , Mauricio Delbracio

Image registration is the inference of transformations relating noisy and distorted images. It is fundamental in computer vision, experimental physics, and medical imaging. Many algorithms and analyses exist for inferring shift, rotation,…

数据分析、统计与概率 · 物理学 2019-02-21 Colin B. Clement , Matthew Bierbaum , James P. Sethna

Face recognition systems are increasingly used in biometric security for convenience and effectiveness. However, they remain vulnerable to spoofing attacks, where attackers use photos, videos, or masks to impersonate legitimate users. This…

计算机视觉与模式识别 · 计算机科学 2024-10-23 Arman Keresh , Pakizar Shamoi

This study evaluates the efficacy of vision transformer models, specifically Swin transformers, in enhancing the diagnostic accuracy of ear diseases compared to traditional convolutional neural networks. With a reported 27% misdiagnosis…

计算机视觉与模式识别 · 计算机科学 2025-11-14 James Ndubuisi , Fernando Auat , Marta Vallejo

In this paper, we present token labeling -- a new training objective for training high-performance vision transformers (ViTs). Different from the standard training objective of ViTs that computes the classification loss on an additional…

计算机视觉与模式识别 · 计算机科学 2021-06-10 Zihang Jiang , Qibin Hou , Li Yuan , Daquan Zhou , Yujun Shi , Xiaojie Jin , Anran Wang , Jiashi Feng

Deep learning models are transforming agricultural applications by enabling automated phenotyping, monitoring, and yield estimation. However, their effectiveness heavily depends on large amounts of annotated training data, which can be…

计算机视觉与模式识别 · 计算机科学 2025-04-11 Rajhans Singh , Rafael Bidese Puhl , Kshitiz Dhakal , Sudhir Sornapudi

With the advent of recent advances in unsupervised learning, efficient training of a deep network for image denoising without pairs of noisy and clean images has become feasible. However, most current unsupervised denoising methods are…

图像与视频处理 · 电气工程与系统科学 2020-12-08 Kanggeun Lee , Won-Ki Jeong

This study demonstrates a cost-effective approach to semantic segmentation using self-supervised vision transformers (SSVT). By freezing the SSVT backbone and training a lightweight segmentation head, our approach effectively utilizes…

计算机视觉与模式识别 · 计算机科学 2024-01-24 Seungho Lee , Seoungyoon Kang , Hyunjung Shim

Vision-language models trained with contrastive learning on paired medical images and reports show strong zero-shot diagnostic capabilities, yet the effect of training batch composition on learned representations remains unexplored for 3D…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Shivika , Kartik Bose , Pankaj Gupta

Vision Transformers (ViT), when paired with large-scale pretraining, have shown remarkable performance across various computer vision tasks, primarily due to their weak inductive bias. However, while such weak inductive bias aids in…

计算机视觉与模式识别 · 计算机科学 2024-06-11 Dongyoon Hwang , Byungkun Lee , Hojoon Lee , Hyunseung Kim , Jaegul Choo

Zero-shot anomaly segmentation using pre-trained foundation models is a promising approach that enables effective algorithms without expensive, domain-specific training or fine-tuning. Ensuring that these methods work across various…

计算机视觉与模式识别 · 计算机科学 2024-05-14 Kevin Stangl , Marius Arvinte , Weilin Xu , Cory Cornelius

Automated sorting is crucial for improving the efficiency and scalability of textile recycling, but accurately identifying material composition and detecting contaminants from sensor data remains challenging. This paper investigates the use…

计算机视觉与模式识别 · 计算机科学 2025-06-10 Yannis Spyridis , Vasileios Argyriou

DC voltage versus current measurements of superconductors in a magnetic field are widely interpreted to imply that a phase transition occurs into a state of zero resistance. We show that the widely-used scaling function approach has a…

超导电性 · 物理学 2015-06-24 D. R. Strachan , M. C. Sullivan , C. J. Lobb

The vision transformer-based foundation models, such as ViT or Dino-V2, are aimed at solving problems with little or no finetuning of features. Using a setting of prototypical networks, we analyse to what extent such foundation models can…

计算机视觉与模式识别 · 计算机科学 2024-02-26 Dmitry Kangin , Plamen Angelov

A variety of recent methods guide large language model outputs via the inference-time addition of steering vectors to residual-stream or attention-head representations. In contrast, we propose to inject steering vectors directly into the…

机器学习 · 计算机科学 2025-09-23 Max Torop , Aria Masoomi , Masih Eskandar , Jennifer Dy

Vision-language models such as CLIP have shown great impact on diverse downstream tasks for zero-shot or label-free predictions. However, when it comes to low-level vision such as image restoration their performance deteriorates…

计算机视觉与模式识别 · 计算机科学 2024-02-29 Ziwei Luo , Fredrik K. Gustafsson , Zheng Zhao , Jens Sjölund , Thomas B. Schön

Zero sample learning is an effective method for data deficiency. The existing embedded zero sample learning methods only use the known classes to construct the embedded space, so there is an overfitting of the known classes in the testing…

计算机视觉与模式识别 · 计算机科学 2024-08-14 Xiaohan Cheng , Taiyuan Mei , Yun Zi , Qi Wang , Zijun Gao , Haowei Yang

Multi-modal foundation models such as CLIP have showcased impressive zero-shot capabilities. However, their applicability in resource-constrained environments is limited due to their large number of parameters and high inference time. While…

计算机视觉与模式识别 · 计算机科学 2024-04-26 Niclas Popp , Jan Hendrik Metzen , Matthias Hein