English
Related papers

Related papers: MSDiagnosis: A Benchmark for Evaluating Large Lang…

200 papers

Diagnostic prediction and clinical reasoning are critical tasks in healthcare applications. While Large Language Models (LLMs) have shown strong capabilities in commonsense reasoning, they still struggle with diagnostic reasoning due to…

Computation and Language · Computer Science 2026-04-28 Yimin Deng , Zhenxi Lin , Yejing Wang , Guoshuai Zhao , Pengyue Jia , Zichuan Fu , Derong Xu , Yefeng Zheng , Xiangyu Zhao , Li Zhu , Xian Wu , Xueming Qian

As opposed to evaluating computation and logic-based reasoning, current benchmarks for evaluating large language models (LLMs) in medicine are primarily focused on question-answering involving domain knowledge and descriptive reasoning.…

Large language models (LLMs) hold promise in clinical decision support but face major challenges in safety evaluation and effectiveness validation. We developed the Clinical Safety-Effectiveness Dual-Track Benchmark (CSEDB), a…

Can the rapid advances in code generation, function calling, and data analysis using large language models (LLMs) help automate the search and verification of hypotheses purely from a set of provided datasets? To evaluate this question, we…

The emergence of groundbreaking large language models capable of performing complex reasoning tasks holds significant promise for addressing various scientific challenges, including those arising in complex clinical scenarios. To enable…

Computation and Language · Computer Science 2025-05-30 Yakun Zhu , Zhongzhen Huang , Linjie Mu , Yutong Huang , Wei Nie , Jiaji Liu , Shaoting Zhang , Pengfei Liu , Xiaofan Zhang

As large language models (LLMs) become increasingly integrated into clinical decision-making, ensuring transparent and trustworthy reasoning is essential. However, existing evaluation strategies of LLMs' medical reasoning capability either…

With the rapid advancement of Multimodal Large Language Models (MLLMs), numerous evaluation benchmarks have emerged. However, comprehensive assessments of their performance across diverse industrial applications remain limited. In this…

Computation and Language · Computer Science 2025-01-29 Dongyi Yi , Guibo Zhu , Chenglin Ding , Zongshu Li , Dong Yi , Jinqiao Wang

The unprecedented performance of large language models (LLMs) requires comprehensive and accurate evaluation. We argue that for LLMs evaluation, benchmarks need to be comprehensive and systematic. To this end, we propose the ZhuJiu…

Computation and Language · Computer Science 2023-08-29 Baoli Zhang , Haining Xie , Pengfan Du , Junhao Chen , Pengfei Cao , Yubo Chen , Shengping Liu , Kang Liu , Jun Zhao

Mental disorders are highly prevalent worldwide, but the shortage of psychiatrists and the inherent subjectivity of interview-based diagnosis create substantial barriers to timely and consistent mental-health assessment. Progress in…

Multiagent Systems · Computer Science 2026-02-12 Shihao Xu , Tiancheng Zhou , Jiatong Ma , Yanli Ding , Yiming Yan , Ming Xiao , Guoyi Li , Haiyang Geng , Yunyun Han , Jianhua Chen , Yafeng Deng

Open-ended medical LLM evaluation remains weakly grounded in physician-calibrated coverage of clinically relevant response criteria, especially in localized clinical settings. We introduce \textsc{ClinConsensus}, a Chinese medical benchmark…

Computation and Language · Computer Science 2026-05-28 Xiang Zheng , Han Li , Wenjie Luo , Weiqi Zhai , Yiyuan Li , Chuanmiao Yan , Xue Yang , Kailuan Wu , Ruyi Xu , Tianyun Lu , Tianyi Tang , Yubo Ma , Kexin Yang , Dayiheng Liu , Sen Yang , Lin Qu , Bing Zhao , Hu Wei

Clinical diagnosis begins with doctor-patient interaction, during which physicians iteratively gather information, determine examination and refine differential diagnosis through patients' response. This dynamic clinical-reasoning process…

Computation and Language · Computer Science 2025-12-30 Yuqi Tang , Jing Yu , Zichang Su , Kehua Feng , Zhihui Zhu , Libin Wang , Lei Liang , Qiang Zhang , Keyan Ding , Huajun Chen

The integration of Artificial Intelligence (AI), especially Large Language Models (LLMs), into the clinical diagnosis process offers significant potential to improve the efficiency and accessibility of medical care. While LLMs have shown…

Computation and Language · Computer Science 2024-10-15 Mingyu Derek Ma , Chenchen Ye , Yu Yan , Xiaoxuan Wang , Peipei Ping , Timothy S Chang , Wei Wang

Large language models (LLMs) have demonstrated strong reasoning abilities across specialized domains, motivating research into their application to legal reasoning. However, existing legal benchmarks often conflate factual recall with…

Artificial Intelligence · Computer Science 2025-11-21 Wenhan Yu , Xinbo Lin , Lanxin Ni , Jinhua Cheng , Lei Sha

Multimodal Large Language Models are primarily trained and evaluated on aligned image-text pairs, which leaves their ability to detect and resolve real-world inconsistencies largely unexplored. In open-domain applications visual and textual…

Developing Large Language Models (LLMs) with robust long-context capabilities has been the recent research focus, resulting in the emergence of long-context LLMs proficient in Chinese. However, the evaluation of these models remains…

Computation and Language · Computer Science 2024-10-17 Zexuan Qiu , Jingjing Li , Shijue Huang , Xiaoqi Jiao , Wanjun Zhong , Irwin King

Large language models (LLMs) have performed remarkably well in various natural language processing tasks by benchmarking, including in the Western medical domain. However, the professional evaluation benchmarks for LLMs have yet to be…

Computation and Language · Computer Science 2024-06-04 Wenjing Yue , Xiaoling Wang , Wei Zhu , Ming Guan , Huanran Zheng , Pengfei Wang , Changzhi Sun , Xin Ma

Large language models (LLMs) have demonstrated significant potential in advancing various fields of research and society. However, the current community of LLMs overly focuses on benchmarks for analyzing specific foundational skills (e.g.…

Curated datasets for healthcare are often limited due to the need of human annotations from experts. In this paper, we present MedEval, a multi-level, multi-task, and multi-domain medical benchmark to facilitate the development of language…

Computation and Language · Computer Science 2023-11-16 Zexue He , Yu Wang , An Yan , Yao Liu , Eric Y. Chang , Amilcare Gentili , Julian McAuley , Chun-Nan Hsu

We introduce FinMMDocR, a novel bilingual multimodal benchmark for evaluating multimodal large language models (MLLMs) on real-world financial numerical reasoning. Compared to existing benchmarks, our work delivers three major advancements.…

Biomedical language understanding benchmarks are the driving forces for artificial intelligence applications with large language model (LLM) back-ends. However, most current benchmarks: (a) are limited to English which makes it challenging…

Computation and Language · Computer Science 2023-10-24 Wei Zhu , Xiaoling Wang , Huanran Zheng , Mosha Chen , Buzhou Tang