中文
相关论文

相关论文: DeVisE: Behavioral Testing of Medical Large Langua…

200 篇论文

Large language models (LLMs) have the potential to transform medicine, but real-world clinical scenarios contain extraneous information that can hinder performance. The rise of assistive technologies like ambient dictation, which…

Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making. However, it remains unclear whether Large Language Models…

计算与语言 · 计算机科学 2025-05-20 Xiaomin Li , Mingye Gao , Yuexing Hao , Taoran Li , Guangya Wan , Zihan Wang , Yijun Wang

Large language models perform well on static medical examinations, yet clinical diagnosis often requires iterative evidence gathering under uncertainty. Building on prior interactive evaluation efforts, we introduce an OSCE-inspired…

人工智能 · 计算机科学 2026-05-22 Chen Zhan , Xihe Qiu , Xiaoyu Tan , Xibing Zhuang , Gengchen Ma , Yue Zhang , Shuo Li , Peifeng Liu , Xiaoxiao Ge , Liang Liu , Lu Gan

Multimodal large language models (MLLMs) have shown great potential in medical applications, yet existing benchmarks inadequately capture real-world clinical complexity. We introduce MEDSYN, a multilingual, multimodal benchmark of highly…

计算与语言 · 计算机科学 2026-04-20 Boqi Chen , Xudong Liu , Jiachuan Peng , Marianne Frey-Marti , Bang Zheng , Kyle Lam , Lin Li , Jianing Qiu

Depression is a widespread mental disorder that affects millions worldwide. While automated depression assessment shows promise, most studies rely on limited or non-clinically validated data, and often prioritize complex model design over…

计算与语言 · 计算机科学 2025-08-07 Zhuang Chen , Guanqun Bi , Wen Zhang , Jiawei Hu , Aoyun Wang , Xiyao Xiao , Kun Feng , Minlie Huang

Objective Structured Clinical Examinations (OSCEs) are widely used to assess medical students' communication skills, but scoring interview-based assessments is time-consuming and potentially subject to human bias. This study explored the…

Large Vision and Language Models have enabled significant advances in fully supervised and zero-shot visual tasks. These large architectures serve as the baseline to what is currently known as Instruction Tuning Large Vision and Language…

计算机视觉与模式识别 · 计算机科学 2025-06-02 Andrés Villa , Juan Carlos León Alcázar , Alvaro Soto , Bernard Ghanem

Objective: Survival analysis is central to medical prediction, yet large language models (LLMs) are rarely used as end-to-end survival models because censoring prevents straightforward supervised fine-tuning. Here we present LLMSurvival, a…

人工智能 · 计算机科学 2026-05-26 Yishu Wei , Hexin Dong , Yi Lin , Jiahe Qian , Yi Liu , Yifan Peng

Robust benchmarks are crucial for evaluating Multimodal Large Language Models (MLLMs). Yet we find that models can ace many multimodal benchmarks without strong visual understanding, instead exploiting biases, linguistic priors, and…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Ellis Brown , Jihan Yang , Shusheng Yang , Rob Fergus , Saining Xie

As Vision Language Models (VLMs) gain widespread use, their fairness remains under-explored. In this paper, we analyze demographic biases across five models and six datasets. We find that portrait datasets like UTKFace and CelebA are the…

计算与语言 · 计算机科学 2025-04-01 Kuleen Sasse , Shan Chen , Jackson Pond , Danielle Bitterman , John Osborne

Despite achieving high accuracy on medical benchmarks, LLMs exhibit the Einstellung Effect in clinical diagnosis--relying on statistical shortcuts rather than patient-specific evidence, causing misdiagnosis in atypical cases. Existing…

计算与语言 · 计算机科学 2026-01-13 Wenting Chen , Zhongrui Zhu , Guolin Huang , Wenxuan Wang

Effective de-escalation is critical for law enforcement safety and community trust, yet traditional training methods lack scalability and realism. While Large Language Models (LLMs) enable dynamic, open-ended simulations, their substantial…

计算与语言 · 计算机科学 2026-05-08 Md Hasebul Hasan , Krity Haque Charu , Eshwara Prasad Sridhar , Shuchisnigdha Deb , Mohammad A. Islam

Reliable deployment of Vision-Language Models (VLMs) in radiology requires validation metrics that go beyond surface-level text similarity to ensure clinical fidelity and demographic fairness. This paper investigates a critical blind spot…

计算与语言 · 计算机科学 2026-03-03 Aditya Parikh , Aasa Feragen , Sneha Das , Stella Frank

Several recent works seek to develop foundation models specifically for medical applications, adapting general-purpose large language models (LLMs) and vision-language models (VLMs) via continued pretraining on publicly available biomedical…

计算与语言 · 计算机科学 2024-11-21 Daniel P. Jeong , Saurabh Garg , Zachary C. Lipton , Michael Oberst

As large language models (LLMs) become an important way of information access, there have been increasing concerns that LLMs may intensify the spread of unethical content, including implicit bias that hurts certain populations without…

计算与语言 · 计算机科学 2025-07-14 Yuchen Wen , Keping Bi , Wei Chen , Jiafeng Guo , Xueqi Cheng

Large language models (LLMs) often match or exceed clinician-level performance on medical benchmarks, yet very few are evaluated on real clinical data or examined beyond headline metrics. We present, to our knowledge, the first evaluation…

Medical vision--language models (VLMs) are usually evaluated on intact image--question pairs, but trustworthy clinical use requires a stronger property: a model must recognise when the evidential basis for an answer has failed. We study…

计算机视觉与模式识别 · 计算机科学 2026-05-25 Hanqi Jiang , Junhao Chen , Mingyu Kang , Hyeokjae Kwon , Yi Pan , Lifeng Chen , Weihang You , Haozhen Gong , Ruiyu Yan , Jinglei Lv , Lin Zhao , Hui Ren , Quanzheng Li , Tianming Liu , Xiang Li

There is increasing interest in the application large language models (LLMs) to the medical field, in part because of their impressive performance on medical exam questions. While promising, exam questions do not reflect the complexity of…

Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, determine examination and refine differential diagnosis through patients' response. This dynamic clinical-reasoning process…

计算与语言 · 计算机科学 2025-12-30 Yuqi Tang , Jing Yu , Zichang Su , Kehua Feng , Zhihui Zhu , Libin Wang , Lei Liang , Qiang Zhang , Keyan Ding , Huajun Chen

Standard single-turn, static benchmarks fall short in evaluating the nuanced capabilities of Large Language Models (LLMs) on complex tasks such as software engineering. In this work, we propose a novel interactive evaluation framework that…