中文
相关论文

相关论文: MedDialBench: Benchmarking LLM Diagnostic Robustne…

200 篇论文

Large language models (LLM) have achieved impressive performance on medical question-answering benchmarks. However, high benchmark accuracy does not imply that the performance generalizes to real-world clinical settings. Medical…

计算与语言 · 计算机科学 2024-09-04 Robert Osazuwa Ness , Katie Matton , Hayden Helm , Sheng Zhang , Junaid Bajwa , Carey E. Priebe , Eric Horvitz

Despite impressive performance on standard benchmarks, multimodal large language models (MLLMs) face critical challenges in real-world clinical environments where medical images inevitably suffer various quality degradations. Existing…

The deployment of large language models (LLMs) as interactive agents has exposed a category of behavioral failure that prevailing terminology, principally hallucination, fails to adequately characterize. This paper introduces LLM Psychosis…

计算机与社会 · 计算机科学 2026-04-30 Ashutosh Raj

Large language models (LLMs) are increasingly used to support question answering and decision-making in high-stakes, domain-specific settings such as natural hazard response and infrastructure planning, where effective answers must convey…

计算与语言 · 计算机科学 2026-02-11 Homaira Huda Shomee , Rochana Chaturvedi , Yangxinyu Xie , Tanwi Mallick

The proliferation of Large Language Models (LLMs) has intensified concerns about manipulative or deceptive behaviors that can undermine user autonomy, trust, and well-being. Existing safety benchmarks predominantly rely on coarse binary…

人工智能 · 计算机科学 2025-12-30 Sadia Asif , Israel Antonio Rosales Laguan , Haris Khan , Shumaila Asif , Muneeb Asif

Large Language Models (LLMs) have demonstrated remarkable capabilities on general text; however, their proficiency in specialized scientific domains that require deep, interconnected knowledge remains largely uncharacterized. Metabolomics…

计算与语言 · 计算机科学 2025-10-17 Yuxing Lu , Xukai Zhao , J. Ben Tamo , Micky C. Nnamdi , Rui Peng , Shuang Zeng , Xingyu Hu , Jinzhuo Wang , May D. Wang

The rapid advancement of large language models (LLMs) has accelerated their integration into clinical decision support, particularly in prescription review. To enable systematic and fine-grained evaluation, we developed RxBench, a…

Recent advancements in reasoning-enhanced large language models (LLMs), such as DeepSeek-R1 and OpenAI-o3, have demonstrated significant progress. However, their application in professional medical contexts remains underexplored,…

计算与语言 · 计算机科学 2025-03-11 Pengcheng Qiu , Chaoyi Wu , Shuyu Liu , Weike Zhao , Zhuoxia Chen , Hongfei Gu , Chuanjin Peng , Ya Zhang , Yanfeng Wang , Weidi Xie

With the increasing use of large language models (LLMs) in medical decision-support, it is essential to evaluate not only their final answers but also the reliability of their reasoning. Two key risks are Chain-of-Thought (CoT) faithfulness…

计算机视觉与模式识别 · 计算机科学 2025-08-25 Kaiyuan Ji , Yijin Guo , Zicheng Zhang , Xiangyang Zhu , Yuan Tian , Ning Liu , Guangtao Zhai

Recent advances in large language models (LLMs) have shown promising results in medical diagnosis, with some studies indicating superior performance compared to human physicians in specific scenarios. However, the diagnostic capabilities of…

人工智能 · 计算机科学 2025-03-24 Zhoujian Sun , Ziyi Liu , Cheng Luo , Jiebin Chu , Zhengxing Huang

Clinical robustness is critical to the safe deployment of medical Large Language Models (LLMs), but key questions remain about how LLMs and humans may differ in response to the real-world variability typified by clinical settings. To…

人工智能 · 计算机科学 2025-06-23 Abinitha Gourabathina , Yuexing Hao , Walter Gerych , Marzyeh Ghassemi

Current evaluation paradigms for large language models (LLMs) represent a critical blind spot in AI research--relying on opaque numerical metrics that conceal fundamental limitations in spatial reasoning while providing no intuitive…

计算与语言 · 计算机科学 2025-11-05 Liuhao Lin , Ke Li , Zihan Xu , Yuchen Shi , Yulei Qin , Yan Zhang , Xing Sun , Rongrong Ji

While LLM agents have demonstrated remarkable task-oriented abilities such as planning, reasoning, and action, few works have treated them as complete human personalities where emotional dimensions hold equal importance. In this paper, we…

计算与语言 · 计算机科学 2026-05-29 Weihan Peng , Chenxu Zhang , Qianao Wang , Yuling Shi , Heng Lian , Qihong Mao , Jiahao Pang , Chunliang Feng , Bowen Li , Xiaodong Gu

With the rapid adoption of LLM-based chatbots, there is a pressing need to evaluate what humans and LLMs can achieve together. However, standard benchmarks, such as MMLU, measure LLM capabilities in isolation (i.e., "AI-alone"). Here, we…

计算与语言 · 计算机科学 2025-08-13 Serina Chang , Ashton Anderson , Jake M. Hofman

Large language models (LLMs) excel at explicit reasoning, but their implicit computational strategies remain underexplored. Decades of psychophysics research show that humans intuitively process and integrate noisy signals using…

计算与语言 · 计算机科学 2025-12-03 Julian Ma , Jun Wang , Zafeirios Fountas

We examine if frontier chat-based large language models (LLMs) adjust their outputs based on neurodivergence (ND) context in system prompts and describe the nature of these adjustments. Specifically, we propose NDBench, a 576-output…

计算与语言 · 计算机科学 2026-05-04 Ishan Gupta , Pavlo Buryi

Patients are increasingly turning to large language models (LLMs) with medical questions that are complex and difficult to articulate clearly. However, LLMs are sensitive to prompt phrasings and can be influenced by the way questions are…

计算与语言 · 计算机科学 2026-04-08 Hye Sun Yun , Geetika Kapoor , Michael Mackert , Ramez Kouzy , Wei Xu , Junyi Jessy Li , Byron C. Wallace

Multimodal large language models (MLLMs) have advanced clinical tasks for common conditions, but their performance on rare diseases remains largely untested. In rare-disease scenarios, clinicians often lack prior clinical knowledge, forcing…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Junzhi Ning , Jiashi Lin , Yingying Fang , Wei Li , Jiyao Liu , Cheng Tang , Chenglong Ma , Wenhao Tang , Tianbin Li , Ziyan Huang , Guang Yang , Junjun He

Artificial intelligence has significantly advanced healthcare, particularly through large language models (LLMs) that excel in medical question answering benchmarks. However, their real-world clinical application remains limited due to the…

计算与语言 · 计算机科学 2024-07-01 Zhihao Fan , Jialong Tang , Wei Chen , Siyuan Wang , Zhongyu Wei , Jun Xi , Fei Huang , Jingren Zhou

Advancements in dialogue systems powered by large language models (LLMs) have outpaced the development of reliable evaluation metrics, particularly for diverse and creative responses. We present a benchmark for evaluating the robustness of…

计算与语言 · 计算机科学 2025-01-14 Justin Vasselli , Adam Nohejl , Taro Watanabe