English
Related papers

Related papers: LogitDynamics: Reliable ViT Error Detection from L…

200 papers

Self-supervised monocular depth estimation is an attractive solution that does not require hard-to-source depth labels for training. Convolutional neural networks (CNNs) have recently achieved great success in this task. However, their…

Computer Vision and Pattern Recognition · Computer Science 2023-03-24 Chaoqiang Zhao , Youmin Zhang , Matteo Poggi , Fabio Tosi , Xianda Guo , Zheng Zhu , Guan Huang , Yang Tang , Stefano Mattoccia

Large Vision Language Models (LVLMs) achieve strong multimodal reasoning but frequently exhibit hallucinations and incorrect responses with high certainty, which hinders their usage in high-stakes domains. Existing verbalized confidence…

Computer Vision and Pattern Recognition · Computer Science 2026-04-13 Wenyi Xiao , Xinchi Xu , Leilei Gan

Vision Transformers (ViTs) are widely adopted in medical imaging tasks, and some existing efforts have been directed towards vision-language training for Chest X-rays (CXRs). However, we envision that there still exists a potential for…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Umar Marikkar , Sara Atito , Muhammad Awais , Adam Mahdi

Features, logits, and labels are the three primary data when a sample passes through a deep neural network. Feature perturbation and label perturbation receive increasing attention in recent years. They have been proven to be useful in…

Machine Learning · Computer Science 2022-09-27 Mengyang Li , Fengguang Su , Ou Wu , Ji Zhang

Causal inference has emerged as a promising approach to mitigate long-tail classification by handling the biases introduced by class imbalance. However, along with the change of advanced backbone models from Convolutional Neural Networks…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Xiaoshuo Yan , Zhaochuan Li , Lei Meng , Zhuang Qi , Wei Wu , Zixuan Li , Xiangxu Meng

Vision-language models (VLMs) have enabled strong zero-shot classification through image-text alignment. Yet, their purely visual inference capabilities remain under-explored. In this work, we conduct a comprehensive evaluation of both…

Computer Vision and Pattern Recognition · Computer Science 2025-09-12 Illia Volkov , Nikita Kisel , Klara Janouskova , Jiri Matas

Vision Transformers (ViTs) are becoming a very popular paradigm for vision tasks as they achieve state-of-the-art performance on image classification. However, although early works implied that this network structure had increased…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Hugo Lemarchant , Liangzi Li , Yiming Qian , Yuta Nakashima , Hajime Nagahara

Large Language Models (LLMs) frequently generate plausible but non-factual content, a phenomenon known as hallucination. While existing detection methods typically rely on computationally expensive sampling-based consistency checks or…

Machine Learning · Computer Science 2026-05-07 Dan Wilson , Mohamed Akrout

The remarkable success of pretrain-then-finetune paradigm has led to a proliferation of available pre-trained models for vision tasks. This surge presents a significant challenge in efficiently choosing the most suitable pre-trained models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Zixuan Hu , Xiaotong Li , Shixiang Tang , Jun Liu , Yichun Hu , Ling-Yu Duan

Image scoring is a crucial task in numerous real-world applications. To trust a model's judgment, understanding its rationale is essential. This paper proposes a novel training method for Vision Language Models (VLMs) to generate not only…

Computer Vision and Pattern Recognition · Computer Science 2025-06-04 Naoto Tanji , Toshihiko Yamasaki

Given the higher information load processed by large vision-language models (LVLMs) compared to single-modal LLMs, detecting LVLM hallucinations requires more human and time expense, and thus rise a wider safety concerns. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Ruiyang Zhang , Hu Zhang , Zhedong Zheng

Vision transformers (ViTs) have achieved promising results on a variety of Computer Vision tasks, however their quadratic complexity in the number of input tokens has limited their application specially in resource-constrained settings.…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Wentao Zhu

Large language models (LLMs) are often confidently wrong, making reliable uncertainty estimation (UE) essential. Output-based heuristics are cheap but brittle, while probing internal representations is effective yet high-dimensional and…

Machine Learning · Computer Science 2026-03-25 Zvi N. Badash , Yonatan Belinkov , Moti Freiman

Large language models are increasingly deployed in settings where reliability matters, yet output-level uncertainty signals such as token probabilities, entropy, and self-consistency can become brittle under calibration--deployment…

Computation and Language · Computer Science 2026-04-20 Yanli Wang , Peng Kuang , Xiaoyu Han , Kaidi Xu , Haohan Wang

Bandwidth constraints during signal acquisition frequently impede real-time detection applications. Hyperspectral data is a notable example, whose vast volume compromises real-time hyperspectral detection. To tackle this hurdle, we…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Lingfeng Liu , Dong Ni , Hangjie Yuan

Vision transformers (ViT) have demonstrated impressive performance across various machine vision problems. These models are based on multi-head self-attention mechanisms that can flexibly attend to a sequence of image patches to encode…

Computer Vision and Pattern Recognition · Computer Science 2021-11-29 Muzammal Naseer , Kanchana Ranasinghe , Salman Khan , Munawar Hayat , Fahad Shahbaz Khan , Ming-Hsuan Yang

Deep networks are currently the state-of-the-art for sensory perception in autonomous driving and robotics. However, deep models often generate overconfident predictions precluding proper probabilistic interpretation which we argue is due…

Machine Learning · Computer Science 2020-08-25 G. Melotti , C. Premebida , J. J. Bird , D. R. Faria , N. Gonçalves

Translucency is prevalent in everyday scenes. As such, perception of transparent objects is essential for robots to perform manipulation. Compared with texture-rich or texture-less Lambertian objects, transparency induces significant…

Robotics · Computer Science 2020-03-24 Zheming Zhou , Xiaotong Chen , Odest Chadwicke Jenkins

Vision Transformers (ViTs) have delivered remarkable progress through global self-attention, yet their quadratic complexity can become prohibitive for high-resolution inputs. In this work, we present ViT-Linearizer, a cross-architecture…

Computer Vision and Pattern Recognition · Computer Science 2026-02-27 Guoyizhe Wei , Rama Chellappa

Large transformer models are known to produce high-norm tokens. In vision transformers (ViTs), such tokens have been mathematically modeled through the singular vectors of the linear approximations of layers. However, in large language…

Computation and Language · Computer Science 2025-07-01 Haoqi Wang , Tong Zhang , Mathieu Salzmann