中文
相关论文

相关论文: Fundus Image-based Visual Acuity Assessment with P…

200 篇论文

Recent convolutional neural networks (CNNs) have led to impressive performance but often suffer from poor calibration. They tend to be overconfident, with the model confidence not always reflecting the underlying true ambiguity and…

机器学习 · 计算机科学 2020-07-14 Beidi Chen , Weiyang Liu , Zhiding Yu , Jan Kautz , Anshumali Shrivastava , Animesh Garg , Anima Anandkumar

Since the early 2000s, computational visual saliency has been a very active research area. Each year, more and more new models are published in the main computer vision conferences. Nowadays, one of the big challenges is to find a way to…

计算机视觉与模式识别 · 计算机科学 2013-07-23 Nicolas Riche , Matthieu Duvinage , Matei Mancas , Bernard Gosselin , Thierry Dutoit

Application of deep neural networks to medical imaging tasks has in some sense become commonplace. Still, a "thorn in the side" of the deep learning movement is the argument that deep networks are prone to overfitting and are thus unable to…

机器学习 · 计算机科学 2021-07-12 Anthony Sicilia , Xingchen Zhao , Anastasia Sosnovskikh , Seong Jae Hwang

Efficient inference in Large Vision-Language Models is constrained by the high cost of processing thousands of visual tokens, yet it remains unclear which tokens and computations can be safely removed. While attention scores are commonly…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Samyak Jha , Junho Kim

Fundus Fluorescein Angiography (FFA) is a critical tool for assessing retinal vascular dynamics and aiding in the diagnosis of eye diseases. However, its invasive nature and less accessibility compared to Color Fundus (CF) images pose…

图像与视频处理 · 电气工程与系统科学 2024-08-28 Weiyi Zhang , Siyu Huang , Jiancheng Yang , Ruoyu Chen , Zongyuan Ge , Yingfeng Zheng , Danli Shi , Mingguang He

We propose a framework for vision-based human pose estimation and motion prediction that gives conformal prediction guarantees for certifiably safe human-robot collaboration. Our framework combines aleatoric uncertainty estimation with OOD…

机器人学 · 计算机科学 2026-05-18 Jakob Thumm , Marian Frei , Tianle Ni , Matthias Althoff , Marco Pavone

The joint interpretation of multi-modal and multi-view fundus images is critical for retinopathy prevention, as different views can show the complete 3D eyeball field and different modalities can provide complementary lesion areas. Compared…

计算机视觉与模式识别 · 计算机科学 2025-04-15 Yonghao Huang , Leiting Chen , Chuan Zhou

When people query Vision-Language Models (VLMs) but cannot see the accompanying visual context (e.g. for blind and low-vision users), augmenting VLM predictions with natural language explanations can signal which model predictions are…

计算与语言 · 计算机科学 2026-04-23 Keyu He , Tejas Srinivasan , Brihi Joshi , Xiang Ren , Jesse Thomason , Swabha Swayamdipta

Cardiac disease evaluation depends on multiple diagnostic modalities: electrocardiogram (ECG) to diagnose abnormal heart rhythms, and imaging modalities such as Magnetic Resonance Imaging (MRI), Computed Tomography (CT) and echocardiography…

信号处理 · 电气工程与系统科学 2024-12-25 Evariste Njomgue Fotso , Buntheng Ly , Hubert Cochet , Maxime Sermesant

Crucial for building trust in deep learning models for critical real-world applications is efficient and theoretically sound uncertainty quantification, a task that continues to be challenging. Useful uncertainty information is expected to…

机器学习 · 计算机科学 2021-10-28 Zhen Lin , Shubhendu Trivedi , Jimeng Sun

Uncertainty estimation has been widely studied in medical image segmentation as a tool to provide reliability, particularly in deep learning approaches. However, previous methods generally lack effective supervision in uncertainty…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Yuzhu Li , An Sui , Fuping Wu , Xiahai Zhuang

Visual Question Answering (VQA) holds great promise for clinical support, particularly in ophthalmology, where retinal fundus photography is essential for diagnosis. However, ophthalmic VQA benchmarks primarily emphasize answer accuracy,…

计算机视觉与模式识别 · 计算机科学 2026-05-22 Xingyue Wang , Bo Liu , Meng Wang , Zhixuan Zhang , Chengcheng Zhu , Huazhu Fu , Jiang Liu

Color fundus photography (CFP) is central to diagnosing and monitoring retinal disease, yet its acquisition variability (e.g., illumination changes) often degrades image quality, which motivates robust enhancement methods. Unpaired…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Xuanzhao Dong , Wenhui Zhu , Yujian Xiong , Xiwen Chen , Hao Wang , Xin Li , Jiajun Cheng , Zhipeng Wang , Shao Tang , Oana Dumitrascu , Yalin Wang

The uncertainty quantification of sensor measurements coupled with deep learning networks is crucial for many robotics systems, especially for safety-critical applications such as self-driving cars. This paper develops an uncertainty…

机器人学 · 计算机科学 2025-06-23 Qiyuan Wu , Mark Campbell

Ambiguity resolution is a key challenge in multimodal machine translation (MMT), where models must genuinely leverage visual input to map an ambiguous expression to its intended meaning. Although prior work has proposed…

计算与语言 · 计算机科学 2026-05-27 Jingheng Pan , Xintong Wang , Longyue Wang , Liang Ding , Weihua Luo , Chris Biemann

Particle Image Velocimetry (PIV) is a widely used technique for flow measurement that traditionally relies on cross-correlation to track the displacement. Recent advances in deep learning-based methods have significantly improved the…

图像与视频处理 · 电气工程与系统科学 2025-07-29 Wei Wang , Jeremiah Hu , Jia Ai , Yong Lee

Validation of prediction uncertainty (PU) is becoming an essential task for modern computational chemistry. Designed to quantify the reliability of predictions in meteorology, the calibration-sharpness (CS) framework is now widely used to…

化学物理 · 物理学 2022-10-03 Pascal Pernot

Uncertainty quantification is essential for assessing the reliability and trustworthiness of modern AI systems. Among existing approaches, verbalized uncertainty, where models express their confidence through natural language, has emerged…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Weihao Xuan , Qingcheng Zeng , Heli Qi , Junjue Wang , Naoto Yokoya

A recently proposed model observer mimics the foveated nature of the human visual system by processing the entire image with varying spatial detail, executing eye movements and scrolling through slices. The model can predict how human…

信号处理 · 电气工程与系统科学 2021-02-11 Miguel A. Lago , Craig K. Abbey , Miguel P. Eckstein

Conformal prediction (CP) gives distribution-free coverage for modern vision and language models, but it is often forced to make a ranking decision from a single unstable nonconformity score. Standard CP uses one realization, while…

机器学习 · 计算机科学 2026-05-25 Jiapeng Zeng , Yogesh Prabhu , Zhanpeng Zeng , Michael A. Newton , Vikas Singh