English
Related papers

Related papers: CONRep: Uncertainty-Aware Vision-Language Report D…

200 papers

Contrastive Language-Image Pre-training (CLIP) has demonstrated outstanding performance in global image understanding and zero-shot transfer through large-scale text-image alignment. However, the core of medical image analysis often lies in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Jiahui Peng , He Yao , Jingwen Li , Yanzhou Su , Sibo Ju , Yujie Lu , Jin Ye , Hongchun Lu , Xue Li , Lincheng Jiang , Min Zhu , Junlong Cheng

In medical reporting, the accuracy of radiological reports, whether generated by humans or machine learning algorithms, is critical. We tackle a new task in this paper: image-conditioned autocorrection of inaccuracies within these reports.…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Arnold Caleb Asiimwe , Dídac Surís , Pranav Rajpurkar , Carl Vondrick

In assistive robotics serving people with disabilities (PWD), accurate place recognition in built environments is crucial to ensure that robots navigate and interact safely within diverse indoor spaces. Language interfaces, particularly…

Computer Vision and Pattern Recognition · Computer Science 2025-01-10 Yifan Xu , Vineet Kamat , Carol Menassa

Colonoscopic polyp diagnosis is pivotal for early colorectal cancer detection, yet traditional automated reporting suffers from inconsistencies and hallucinations due to the scarcity of high-quality multimodal medical data. To bridge this…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Tianyu Zhou , Junyi Tang , Zehui Li , Dahong Qian , Suncheng Xiang

Automatic radiology report generation holds significant potential to streamline the labor-intensive process of report writing by radiologists, particularly for 3D radiographs such as CT scans. While CT scans are critical for clinical…

Computer Vision and Pattern Recognition · Computer Science 2025-02-26 Che Liu , Zhongwei Wan , Yuqi Wang , Hui Shen , Haozhe Wang , Kangyu Zheng , Mi Zhang , Rossella Arcucci

Computed Tomography Report Generation (CTRG) aims to automate the clinical radiology reporting process, thereby reducing the workload of report writing and facilitating patient care. While deep learning approaches have achieved remarkable…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Hong Liu , Dong Wei , Qiong Peng , Yawen Huang , Xian Wu , Yefeng Zheng , Liansheng Wang

This study addresses the critical challenge of hallucination mitigation in Large Vision-Language Models (LVLMs) for Visual Question Answering (VQA) tasks through a Split Conformal Prediction (SCP) framework. While LVLMs excel in multi-modal…

Computation and Language · Computer Science 2025-05-16 Yuanchang Ye , Weiyan Wen

Electronic health record (EHR) foundation models have been an area ripe for exploration with their improved performance in various medical tasks. Despite the rapid advances, there exists a fundamental limitation: Processing unseen medical…

Artificial Intelligence · Computer Science 2025-08-15 Junmo Kim , Namkyeong Lee , Jiwon Kim , Kwangsoo Kim

Image-to-text radiology report generation aims to automatically produce radiology reports that describe the findings in medical images. Most existing methods focus solely on the image data, disregarding the other patient information…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Nurbanu Aksoy , Serge Sharoff , Selcuk Baser , Nishant Ravikumar , Alejandro F Frangi

Recent work shows that text-only reinforcement learning with verifiable rewards (RLVR) can match or outperform image-text RLVR on multimodal medical VQA benchmarks, suggesting current evaluation protocols may fail to measure causal visual…

Computer Vision and Pattern Recognition · Computer Science 2026-03-05 Anas Zafar , Leema Krishna Murali , Ashish Vashist

Existing research on Retrieval-Augmented Generation (RAG) primarily focuses on improving overall question-answering accuracy, often overlooking the quality of sub-claims within generated responses. Recent methods that attempt to improve RAG…

Information Retrieval · Computer Science 2025-06-27 Naihe Feng , Yi Sui , Shiyi Hou , Jesse C. Cresswell , Ga Wu

Vision-language models (VLMs) often produce chain-of-thought (CoT) explanations that sound plausible yet fail to reflect the underlying decision process, undermining trust in high-stakes clinical use. Existing evaluations rarely catch this…

Multilingual vision-language models have made significant strides in image captioning, yet they still lag behind their English counterparts due to limited multilingual training data and costly large-scale model parameterization.…

Computation and Language · Computer Science 2025-07-29 George Ibrahim , Rita Ramos , Yova Kementchedjhieva

Chest X-ray report generation aims to reduce radiologists' workload by automatically producing high-quality preliminary reports. A critical yet underexplored aspect of this task is the effective use of patient-specific prior knowledge --…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Kang Liu , Zhuoqi Ma , Zikang Fang , Yunan Li , Kun Xie , Qiguang Miao

The proliferation of Vision-Language Models (VLMs) in the past several years calls for rigorous and comprehensive evaluation methods and benchmarks. This work analyzes existing VLM evaluation techniques, including automated metrics,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Alexis Roger , Prateek Humane , Daniel Z. Kaplan , Kshitij Gupta , Qi Sun , George Adamopoulos , Jonathan Siu Chi Lim , Quentin Anthony , Edwin Fennell , Irina Rish

Object-aware reasoning in vision-language tasks poses significant challenges for current models, particularly in handling unseen objects, reducing hallucinations, and capturing fine-grained relationships in complex visual scenes. To address…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Antonio Carlos Rivera , Anthony Moore , Steven Robinson

Continual Learning (CL) is essential for enabling self-evolving large language models (LLMs) to adapt and remain effective amid rapid knowledge growth. Yet, despite its importance, little attention has been given to establishing statistical…

Machine Learning · Computer Science 2025-10-29 Xiaofan Zhou , Lu Cheng

Large Reasoning Models (LRMs) have recently demonstrated significant improvements in complex reasoning. While quantifying generation uncertainty in LRMs is crucial, traditional methods are often insufficient because they do not provide…

Artificial Intelligence · Computer Science 2026-04-16 Yangyi Li , Chenxu Zhao , Mengdi Huai

Vision-language models (VLMs) have shown strong performance in video anomaly detection (VAD) while providing interpretable predictions. However, existing VLM-based VAD methods suffer from a fundamental mismatch between training and…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Darryl Cherian Jacob , Xinyu Liu , Kai Wang , Pan He

The growing integration of vision-language models (VLMs) in medical applications offers promising support for diagnostic reasoning. However, current medical VLMs often face limitations in generalization, transparency, and computational…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Tan-Hanh Pham , Chris Ngo