中文
相关论文

相关论文: Benchmarking Large Language Models on Answering an…

200 篇论文

Reliable causal inference is essential for making decisions in high-stakes areas like medicine, economics, and public policy. However, it remains unclear whether large language models (LLMs) can handle rigorous and trustworthy statistical…

人工智能 · 计算机科学 2026-05-13 Jin Du , Li Chen , Xun Xian , An Luo , Fangqiao Tian , Ganghua Wang , Charles Doss , Xiaotong Shen , Jie Ding

As ChatGPT and GPT-4 spearhead the development of Large Language Models (LLMs), more researchers are investigating their performance across various tasks. But more research needs to be done on the interpretability capabilities of LLMs, that…

计算与语言 · 计算机科学 2023-10-27 Dongfang Li , Jindi Yu , Baotian Hu , Zhenran Xu , Min Zhang

We introduce a novel question-answering (QA) dataset using echocardiogram reports sourced from the Medical Information Mart for Intensive Care database. This dataset is specifically designed to enhance QA systems in cardiology, consisting…

人工智能 · 计算机科学 2025-03-07 Lama Moukheiber , Mira Moukheiber , Dana Moukheiiber , Jae-Woo Ju , Hyung-Chul Lee

We present a refined approach to biomedical question-answering (QA) services by integrating large language models (LLMs) with Multi-BERT configurations. By enhancing the ability to process and prioritize vast amounts of complex biomedical…

计算与语言 · 计算机科学 2024-10-18 Cheng Qian , Xianglong Shi , Shanshan Yao , Yichen Liu , Fengming Zhou , Zishu Zhang , Junaid Akram , Ali Braytee , Ali Anaissi

While large language models (LLMs) achieve near-perfect scores on medical licensing exams, these evaluations inadequately reflect the complexity and diversity of real-world clinical practice. We introduce MedHELM, an extensible evaluation…

Large Language Models (LLMs) have gained significant attention in the medical domain for their human-level capabilities, leading to increased efforts to explore their potential in various healthcare applications. However, despite such a…

Online medical forums have long served as vital platforms where patients seek professional healthcare advice, generating vast amounts of valuable knowledge. However, the informal nature and linguistic complexity of forum interactions pose…

计算与语言 · 计算机科学 2025-10-22 Antonio Romano , Giuseppe Riccio , Mariano Barone , Marco Postiglione , Vincenzo Moscato

Large Language Models (LLMs) have demonstrated strong performance across a wide range of tasks, yet they still struggle with complex mathematical reasoning, a challenge fundamentally rooted in deep structural dependencies. To address this…

人工智能 · 计算机科学 2025-12-01 Lei Zan , Keli Zhang , Ruichu Cai , Lujia Pan

The growing capabilities of Large Language Models (LLMs) show significant potential to enhance healthcare by assisting medical researchers and physicians. However, their reliance on static training data is a major risk when medical…

计算与语言 · 计算机科学 2025-09-05 Juraj Vladika , Mahdi Dhaini , Florian Matthes

Medical multiple-choice question answering (MCQA) is particularly difficult. Questions may describe patient symptoms and ask for the correct diagnosis, which requires domain knowledge and complex reasoning. Standard language modeling…

计算与语言 · 计算机科学 2023-03-14 Damien Sileo , Kanimozhi Uma , Marie-Francine Moens

In response to the pressing need for advanced clinical problem-solving tools in healthcare, we introduce BooksMed, a novel framework based on a Large Language Model (LLM). BooksMed uniquely emulates human cognitive processes to deliver…

Large language models (LLMs) show increasing potential in education, yet benchmarks for non-English languages in specialized domains remain scarce. We introduce MedBench-IT, the first comprehensive benchmark for evaluating LLMs on Italian…

计算与语言 · 计算机科学 2025-09-10 Ruggero Marino Lazzaroni , Alessandro Angioi , Michelangelo Puliga , Davide Sanna , Roberto Marras

Background: The rapid integration of foundation models into clinical practice and public health necessitates a rigorous evaluation of their true clinical reasoning capabilities beyond narrow examination success. Current benchmarks,…

计算机视觉与模式识别 · 计算机科学 2025-12-30 Dingyu Wang , Zimu Yuan , Jiajun Liu , Shanggui Liu , Nan Zhou , Tianxing Xu , Di Huang , Dong Jiang

Real-world clinical practice demands multi-image comparative reasoning, yet current medical benchmarks remain limited to single-frame interpretation. We present MedFrameQA, the first benchmark explicitly designed to test multi-image medical…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Suhao Yu , Haojin Wang , Juncheng Wu , Luyang Luo , Jingshen Wang , Cihang Xie , Pranav Rajpurkar , Carl Yang , Yang Yang , Kang Wang , Yannan Yu , Yuyin Zhou

Ensuring the general efficacy and goodness for human beings from medical large language models (LLM) before real-world deployment is crucial. However, a widely accepted and accessible evaluation process for medical LLM, especially in the…

Clinical Decision Support Systems (CDSSs) provide reasoning and inquiry guidance for physicians, yet they face notable challenges, including high maintenance costs and low generalization capability. Recently, Large Language Models (LLMs)…

计算与语言 · 计算机科学 2026-04-24 Yue Guo , Fanfu Wang , Jianwei Lv , Xincheng Shi , Yuchen Li , Youya Wang , Yunsheng Zeng , Yujing Liu , Yunhao Qiao , Gen Li , Junfeng Wang , Bo Yuan

Recent advances in large language models (LLMs) have shown impressive ability in biomedical question-answering, but have not been adequately investigated for more specific biomedical applications. This study investigates the performance of…

计算与语言 · 计算机科学 2024-03-29 Shan Chen , Yingya Li , Sheng Lu , Hoang Van , Hugo JWL Aerts , Guergana K. Savova , Danielle S. Bitterman

"Citizen queries" are questions asked by an individual about government policies, guidance, and services that are relevant to their circumstances, encompassing a range of topics including benefits, taxes, immigration, employment, public…

计算机与社会 · 计算机科学 2026-02-05 Neil Majithia , Rajat Shinde , Zo Chapman , Prajun Trital , Jordan Decker , Manil Maskey , Elena Simperl , Nigel Shadbolt

Medical conversational AI (AI) plays a pivotal role in the development of safer and more effective medical dialogue systems. However, existing benchmarks and evaluation frameworks for assessing the information-gathering and diagnostic…

计算与语言 · 计算机科学 2026-01-08 Lecheng Gong , Weimin Fang , Ting Yang , Dongjie Tao , Chunxiao Guo , Peng Wei , Bo Xie , Jinqun Guan , Zixiao Chen , Fang Shi , Jinjie Gu , Junwei Liu

Large language models (LLMs) have been widely adopted in various downstream task domains. However, their abilities to directly recall and apply factual medical knowledge remains under-explored. Most existing medical QA benchmarks assess…

计算与语言 · 计算机科学 2025-08-20 Jiaxi Li , Yiwei Wang , Kai Zhang , Yujun Cai , Bryan Hooi , Nanyun Peng , Kai-Wei Chang , Jin Lu