中文
相关论文

相关论文: MedDialogRubrics: A Comprehensive Benchmark and Ev…

200 篇论文

Evaluating large language models (LLMs) for medical applications remains challenging due to benchmark saturation, limited data accessibility, and insufficient coverage of relevant tasks. Existing suites have either saturated, heavily depend…

Evaluating large language models (LLM) in clinical scenarios is crucial to assessing their potential clinical utility. Existing benchmarks rely heavily on static question-answering, which does not accurately depict the complex, sequential…

人机交互 · 计算机科学 2025-05-27 Samuel Schmidgall , Rojin Ziaei , Carl Harris , Eduardo Reis , Jeffrey Jopling , Michael Moor

We present MedPI, a high-dimensional benchmark for evaluating large language models (LLMs) in patient-clinician conversations. Unlike single-turn question-answer (QA) benchmarks, MedPI evaluates the medical dialogue across 105 dimensions…

计算与语言 · 计算机科学 2026-01-09 Diego Fajardo V. , Oleksii Proniakin , Victoria-Elisabeth Gruber , Razvan Marinescu

Large language models (LLMs) excel at solving problems with clear and complete statements, but often struggle with nuanced environments or interactive tasks which are common in most real-world scenarios. This highlights the critical need…

The development of large language models (LLMs) has brought unprecedented possibilities for artificial intelligence (AI) based medical diagnosis. However, the application perspective of LLMs in real diagnostic scenarios is still unclear…

计算与语言 · 计算机科学 2024-05-21 Zhoujian Sun , Cheng Luo , Ziyi Liu , Zhengxing Huang

Large Language Models (LLMs) have demonstrated significant promise for various applications in healthcare. However, their efficacy in the Arabic medical domain remains unexplored due to the lack of high-quality domain-specific datasets and…

计算与语言 · 计算机科学 2025-08-25 Mouath Abu Daoud , Chaimae Abouzahir , Leen Kharouf , Walid Al-Eisawi , Nizar Habash , Farah E. Shamout

The rise of large language models (LLMs) has transformed healthcare by offering clinical guidance, yet their direct deployment to patients poses safety risks due to limited domain expertise. To mitigate this, we propose repositioning LLMs…

计算与语言 · 计算机科学 2025-10-14 Wenya Xie , Qingying Xiao , Yu Zheng , Xidong Wang , Junying Chen , Ke Ji , Anningzhe Gao , Prayag Tiwari , Xiang Wan , Feng Jiang , Benyou Wang

This paper introduces a framework for the automated evaluation of natural language texts. A manually constructed rubric describes how to assess multiple dimensions of interest. To evaluate a text, a large language model (LLM) is prompted…

计算与语言 · 计算机科学 2025-01-03 Helia Hashemi , Jason Eisner , Corby Rosset , Benjamin Van Durme , Chris Kedzie

Recent advancements in Large Language Models (LLMs) have shown outstanding potential for role-playing applications. Evaluating these capabilities is becoming crucial yet remains challenging. Existing benchmarks mostly adopt a…

计算与语言 · 计算机科学 2025-10-24 Hao Xiang , Tianyi Tang , Yang Su , Bowen Yu , An Yang , Fei Huang , Yichang Zhang , Yaojie Lu , Hongyu Lin , Xianpei Han , Jingren Zhou , Junyang Lin , Le Sun

Current evaluations of medical consultation agents often prioritize outcome-oriented tasks, frequently overlooking the end-to-end process integrity and clinical safety essential for real-world practice. While recent interactive benchmarks…

人工智能 · 计算机科学 2026-01-21 Chuhan Qiao , Jianghua Huang , Daxing Zhao , Ziding Liu , Yanjun Shen , Bing Cheng , Wei Lin , Kai Wu

The primary aim of this research was to address the limitations observed in the medical knowledge of prevalent large language models (LLMs) such as ChatGPT, by creating a specialized language model with enhanced accuracy in medical advice.…

计算与语言 · 计算机科学 2023-06-27 Yunxiang Li , Zihan Li , Kai Zhang , Ruilong Dan , Steve Jiang , You Zhang

Current medical AI systems often fail to replicate real-world clinical reasoning, as they are predominantly trained and evaluated on static text and question-answer tasks. These tuning methods and benchmarks overlook critical aspects like…

计算与语言 · 计算机科学 2026-02-24 Zijie Liu , Xinyu Zhao , Jie Peng , Zhuangdi Zhu , Qingyu Chen , Kaidi Xu , Xia Hu , Tianlong Chen

Large Language Models (LLMs) are transforming artificial intelligence, evolving into task-oriented systems capable of autonomous planning and execution. One of the primary applications of LLMs is conversational AI systems, which must…

计算与语言 · 计算机科学 2025-01-22 Elad Levi , Ilan Kadar

Automatic evaluation is an integral aspect of dialogue system research. The traditional reference-based NLG metrics are generally found to be unsuitable for dialogue assessment. Consequently, recent studies have suggested various unique,…

计算与语言 · 计算机科学 2024-01-23 Chen Zhang , Luis Fernando D'Haro , Yiming Chen , Malu Zhang , Haizhou Li

Despite strong performance in medical question-answering, the clinical adoption of Large Language Models (LLMs) is critically hampered by their opaque 'black-box' reasoning, limiting clinician trust. This challenge is compounded by the…

计算机视觉与模式识别 · 计算机科学 2025-05-13 Chao Ding , Mouxiao Bian , Pengcheng Chen , Hongliang Zhang , Tianbin Li , Lihao Liu , Jiayuan Chen , Zhuoran Li , Yabei Zhong , Yongqi Liu , Haiqing Huang , Dongming Shan , Junjun He , Jie Xu

Interactive medical dialogue benchmarks have shown that LLM diagnostic accuracy degrades significantly when interacting with non-cooperative patients, yet existing approaches either apply adversarial behaviors without graded severity or…

计算与语言 · 计算机科学 2026-04-09 Xiaotian Luo , Xun Jiang , Jiangcheng Wu

The advent of Large Language Models (LLMs) has drastically enhanced dialogue systems. However, comprehensively evaluating the dialogue abilities of LLMs remains a challenge. Previous benchmarks have primarily focused on single-turn…

计算与语言 · 计算机科学 2024-11-06 Ge Bai , Jie Liu , Xingyuan Bu , Yancheng He , Jiaheng Liu , Zhanhui Zhou , Zhuoran Lin , Wenbo Su , Tiezheng Ge , Bo Zheng , Wanli Ouyang

Evaluating large language models (LLMs) in medicine is crucial because medical applications require high accuracy with little room for error. Current medical benchmarks have three main types: medical exam-based, comprehensive medical, and…

Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making. However, it remains unclear whether Large Language Models…

计算与语言 · 计算机科学 2025-05-20 Xiaomin Li , Mingye Gao , Yuexing Hao , Taoran Li , Guangya Wan , Zihan Wang , Yijun Wang

The rapid advancements in large language models (LLMs) have opened up new opportunities for transforming patient engagement in healthcare through conversational AI. This paper presents an overview of the current landscape of LLMs in…

人工智能 · 计算机科学 2024-11-27 Bo Wen , Raquel Norel , Julia Liu , Thaddeus Stappenbeck , Farhana Zulkernine , Huamin Chen