English
Related papers

Related papers: Strong Reasoning Isn't Enough: Evaluating Evidence…

200 papers

Large language models perform well on static medical examinations, yet clinical diagnosis often requires iterative evidence gathering under uncertainty. Building on prior interactive evaluation efforts, we introduce an OSCE-inspired…

Artificial Intelligence · Computer Science 2026-05-22 Chen Zhan , Xihe Qiu , Xiaoyu Tan , Xibing Zhuang , Gengchen Ma , Yue Zhang , Shuo Li , Peifeng Liu , Xiaoxiao Ge , Liang Liu , Lu Gan

Explanations for AI models in high-stakes domains like medicine often lack verifiability, which can hinder trust. To address this, we propose an interactive agent that produces explanations through an auditable sequence of actions. The…

Artificial Intelligence · Computer Science 2025-11-04 Yuhang Huang , Zekai Lin , Fan Zhong , Lei Liu

Disease screening is critical for early detection and timely intervention in clinical practice. However, most current screening models for medical images suffer from limited interpretability and suboptimal performance. They often lack…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Chenyu Lian , Hong-Yu Zhou , Jing Qin

The automation of the medical evidence acquisition and diagnosis process has recently attracted increasing attention in order to reduce the workload of doctors and democratize access to medical care. However, most works proposed in the…

Computation and Language · Computer Science 2022-10-14 Arsene Fansi Tchango , Rishab Goel , Julien Martel , Zhi Wen , Gaetan Marceau Caron , Joumana Ghosn

Recent advances in large language models (LLMs) have enabled promising progress in diagnosis prediction from electronic health records (EHRs). However, existing LLM-based approaches tend to overfit to historically observed diagnoses, often…

Computation and Language · Computer Science 2026-04-14 Hengyu Zhang , Xuyun Zhang , Pengxiang Zhan , Linhao Luo , Hang Lv , Yanchao Tan , Shirui Pan , Carl Yang

Existing reasoning evaluation paradigms suffer from different limitations: fixed benchmarks are increasingly saturated and vulnerable to contamination, while preference-based evaluations rely on subjective judgments. We argue that a core…

Artificial Intelligence · Computer Science 2026-05-19 Baoqing Yue , Zihan Zhu , Yutong Han , Brian Fan , Qian Sun , Jichen Feng , Hufei Yang , Yifan Zhang , Mengdi Wang

Healthcare professionals need effective ways to use, understand, and validate AI-driven clinical decision support systems. Existing systems face two key limitations: complex visualizations and a lack of grounding in scientific evidence. We…

Human-Computer Interaction · Computer Science 2025-07-08 Reza Samimi , Aditya Bhattacharya , Lucija Gosak , Gregor Stiglic , Katrien Verbert

Large language models with reasoning capabilities have demonstrated impressive performance across a wide range of domains. In clinical applications, a transparent, step-by-step reasoning process provides physicians with strong evidence to…

Artificial Intelligence · Computer Science 2025-12-16 Linjie Mu , Yannian Gu , Zhongzhen Huang , Yakun Zhu , Shaoting Zhang , Xiaofan Zhang

Medical diagnostics is a high-stakes and complex domain that is critical to patient care. However, current evaluations of large language models (LLMs) remain limited in capturing key challenges of clinical diagnostic scenarios. Most rely on…

Computation and Language · Computer Science 2026-04-21 Xiangxu Zhang , Lei Li , Yanyun Zhou , Xiao Zhou , Yingying Zhang , Xian Wu

Most automated electronic medical record (EMR) pipelines remain output-oriented: they transcribe, extract, and summarize after the consultation, but they do not explicitly model what is already known, what is still missing, which…

Artificial Intelligence · Computer Science 2026-03-19 Zhenhai Pan , Yan Liu , Jia You

Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnostic investigations. In contrast, current medical benchmarks…

Machine Learning · Computer Science 2025-10-14 Christopher Chiu , Silviu Pitis , Mihaela van der Schaar

Long-term conversational memory requires retrieving evidence scattered across multiple sessions, yet single-pass retrieval fails on temporal and multi-hop questions. Existing iterative methods refine queries via generated content or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-01 Yuyang Li , Yime He , Zeyu Zhang , Dong Gong

Access to the right evidence does not guarantee that large language models (LLMs) will reason with it correctly. This gap between retrieval and reasoning is especially concerning in clinical settings, where outputs must align with…

Artificial Intelligence · Computer Science 2025-11-21 Deepthi Potluri , Aby Mammen Mathew , Jeffrey B DeWitt , Alexander L. Rasgon , Yide Hao , Junyuan Hong , Ying Ding

Current medical retrieval benchmarks primarily emphasize lexical or shallow semantic similarity, overlooking the reasoning-intensive demands that are central to clinical decision-making. In practice, physicians often retrieve authoritative…

Information Retrieval · Computer Science 2026-04-07 Xiangxu Zhang , Lei Li , Xiao Zhou , Zheng Liu

Deep Research (DR) is an emerging agent application that leverages large language models (LLMs) to address open-ended queries. It requires the integration of several capabilities, including multi-step reasoning, cross-document synthesis,…

One of the fundamental cognitive abilities of humans is to quickly resolve uncertainty by generating hypotheses and testing them via active trials. Encountering a novel phenomenon accompanied by ambiguous cause-effect relationships, humans…

Artificial Intelligence · Computer Science 2023-10-31 Manjie Xu , Guangyuan Jiang , Wei Liang , Chi Zhang , Yixin Zhu

Psychiatric questionnaires are highly context sensitive and often only weakly predict subsequent symptom severity, which makes the prognostic relationship difficult to learn. Although flexible nonlinear models can improve predictive…

Machine Learning · Computer Science 2026-03-02 Eric V. Strobl

Rare diseases affect hundreds of millions worldwide, yet diagnosis often spans years. Convectional pipelines decouple noisy evidence extraction from downstream inferential diagnosis, and general/medical large language models (LLMs) face…

While existing benchmarks probe the reasoning abilities of large language models (LLMs) across diverse domains, they predominantly assess passive reasoning, providing models with all the information needed to reach a solution. By contrast,…

Machine Learning · Computer Science 2025-06-11 Zhanke Zhou , Xiao Feng , Zhaocheng Zhu , Jiangchao Yao , Sanmi Koyejo , Bo Han

Interactive medical questioning is essential in real-world clinical consultations, where physicians must actively gather information from patients. While medical Large Language Models (LLMs) have shown impressive capabilities in static…

Computation and Language · Computer Science 2025-08-20 Hongxin Ding , Baixiang Huang , Yue Fang , Weibin Liao , Xinke Jiang , Zheng Li , Junfeng Zhao , Yasha Wang
‹ Prev 1 2 3 10 Next ›