English
Related papers

Related papers: Evaluating AI systems under uncertain ground truth…

200 papers

Artificial intelligence (AI) has demonstrated strong potential in clinical diagnostics, often achieving accuracy comparable to or exceeding that of human experts. A key challenge, however, is that AI reasoning frequently diverges from…

Artificial Intelligence · Computer Science 2026-05-25 Belona Sonna , Alban Grastien

Appropriate reliance is critical to achieving synergistic human-AI collaboration. For instance, when users over-rely on AI assistance, their human-AI team performance is bounded by the model's capability. This work studies how the…

Human-Computer Interaction · Computer Science 2024-01-17 Shiye Cao , Anqi Liu , Chien-Ming Huang

Generative Artificial Intelligence (GenAI) is now widespread in education, yet the efficacy of GenAI systems remains constrained by the quality and interpretation of the labeled data used to train and evaluate them. Studies commonly report…

Computers and Society · Computer Science 2026-04-01 Danielle R. Thomas , Conrad Borchers , Kirk P. Vanacore , Kenneth R. Koedinger , René F. Kizilcec

Fairness has become increasingly pivotal in medical image recognition. However, without mitigating bias, deploying unfair medical AI systems could harm the interests of underprivileged populations. In this paper, we observe that while…

Computer Vision and Pattern Recognition · Computer Science 2024-02-21 Ching-Hao Chiu , Hao-Wei Chung , Yu-Jen Chen , Yiyu Shi , Tsung-Yi Ho

Artificial intelligence (AI) is increasingly integrated into modern healthcare, offering powerful support for clinical decision-making. However, in real-world settings, AI systems may experience performance degradation over time, due to…

Artificial Intelligence · Computer Science 2026-02-05 Hao Guan , David Bates , Li Zhou

Medical image segmentation supports clinical workflows by precisely delineating anatomical structures and lesions. However, medical image datasets medical image datasets suffer from acquisition noise and annotation ambiguity, causing…

Artificial Intelligence · Computer Science 2026-04-14 Ruiyang Li , Fang Liu , Licheng Jiao , Xinglin Xie , Jiayao Hao , Shuo Li , Xu Liu , Jingyi Yang , Lingling Li , Puhua Chen , Wenping Ma

The clinical translation of dermatological AI is hindered by opaque reasoning and systematic performance disparities across skin tones. Here we present SkinGPT-R1, a multimodal large language model that integrates chain-of-thought…

Computer Vision and Pattern Recognition · Computer Science 2026-02-19 Yuhao Shen , Zhangtianyi Chen , Yuanhao He , Yan Xu , Shuping Zhang , Liyuan Sun , Zijian Wang , Yinghao Zhu , Yuyuan Yang , Jiahe Qian , Ziwen Wang , Xinyuan Zhang , Wenbin Liu , Zongyuan Ge , Tao Lu , Siyuan Yan , Juexiao Zhou

Deep Learning approaches in dermatological image classification have shown promising results, yet the field faces significant methodological challenges that impede proper evaluation. This paper presents a dual contribution: first, a…

Image and Video Processing · Electrical Eng. & Systems 2025-02-05 Łukasz Miętkiewicz , Leon Ciechanowski , Dariusz Jemielniak

Recent significant increases in affordable and accessible computational power and data storage have enabled machine learning to provide almost unbelievable classification and prediction performances compared to well-trained humans. There…

Computers and Society · Computer Science 2020-07-28 Gari D. Clifford

Deep neural networks has been increasingly applied in fault diagnostics, where it uses historical data to capture systems behavior, bypassing the need for high-fidelity physical models. However, despite their competence in prediction tasks,…

Machine Learning · Computer Science 2025-09-24 Arman Mohammadi , Mattias Krysander , Daniel Jung , Erik Frisk

Background: In medical imaging, prior studies have demonstrated disparate AI performance by race, yet there is no known correlation for race on medical imaging that would be obvious to the human expert interpreting the images. Methods:…

Clinical dataset labels are rarely certain as annotators disagree and confidence is not uniform across cases. Typical aggregation procedures, such as majority voting, obscure this variability. In simple experiments on medical imaging…

Artificial intelligence (AI) is rapidly advancing in healthcare, enhancing the efficiency and effectiveness of services across various specialties, including cardiology, ophthalmology, dermatology, emergency medicine, etc. AI applications…

Deep neural networks have demonstrated promising performance on image recognition tasks. However, they may heavily rely on confounding factors, using irrelevant artifacts or bias within the dataset as the cue to improve performance. When a…

Computer Vision and Pattern Recognition · Computer Science 2023-03-03 Siyuan Yan , Zhen Yu , Xuelin Zhang , Dwarikanath Mahapatra , Shekhar S. Chandra , Monika Janda , Peter Soyer , Zongyuan Ge

Performance uncertainty quantification is essential for reliable validation and eventual clinical translation of medical imaging artificial intelligence (AI). Confidence intervals (CIs) play a central role in this process by indicating how…

Large language models are increasingly being assembled into medical multi-agent systems that emulate multidisciplinary consultation through specialist roles, peer review and consensus formation. In clinical decision support, however,…

Computation and Language · Computer Science 2026-05-28 Yinghao Zhu , Lei Gu , Zixiang Wang , Haoran Sang , Dehao Sui , Wen Tang , Lan Mi , Yasha Wang , Junyi Gao , Liang Yao , Tianfan Fu , Ewen Harrison , Lequan Yu , Liantao Ma

Fairness and accountability are two essential pillars for trustworthy Artificial Intelligence (AI) in healthcare. However, the existing AI model may be biased in its decision marking. To tackle this issue, we propose an adversarial…

Computer Vision and Pattern Recognition · Computer Science 2021-05-12 Xiaoxiao Li , Ziteng Cui , Yifan Wu , Lin Gu , Tatsuya Harada

In a data-scarce field such as healthcare, where models often deliver predictions on patients with rare conditions, the ability to measure the uncertainty of a model's prediction could potentially lead to improved effectiveness of decision…

Machine Learning · Statistics 2020-05-26 Lotta Meijerink , Giovanni Cinà , Michele Tonutti

Large language models (LLMs) have shown considerable potential in supporting medical diagnosis. However, their effective integration into clinical workflows is hindered by physicians' difficulties in perceiving and trusting LLM…

Human-Computer Interaction · Computer Science 2026-01-28 Yuansong Xu , Yichao Zhu , Haokai Wang , Yuchen Wu , Yang Ouyang , Hanlu Li , Wenzhe Zhou , Xinyu Liu , Chang Jiang , Quan Li

Large language models have demonstrated remarkable performance in a wide range of medical benchmarks. Yet underneath the seemingly promising results lie salient growth areas, especially in cutting-edge frontiers such as multimodal…