English
Related papers

Related papers: Measuring Competency, Not Performance: Item-Aware …

200 papers

Aligning test items to content standards is a critical step in test development to collect validity evidence based on content. Item alignment has typically been conducted by human experts. This judgmental process can be subjective and…

Computation and Language · Computer Science 2025-10-14 Yanbin Fu , Hong Jiao , Tianyi Zhou , Nan Zhang , Ming Li , Qingshu Xu , Sydney Peters , Robert W. Lissitz

As large language models (LLMs) are increasingly integrated into daily life, in roles ranging from high-stakes decision support to companionship, understanding their behavioral dispositions becomes critical. A growing literature uses…

Artificial Intelligence · Computer Science 2026-04-22 Valentin Kriegmair , Dirk U. Wulff

Clinical guidelines, typically structured as decision trees, are central to evidence-based medical practice and critical for ensuring safe and accurate diagnostic decision-making. However, it remains unclear whether Large Language Models…

Computation and Language · Computer Science 2025-05-20 Xiaomin Li , Mingye Gao , Yuexing Hao , Taoran Li , Guangya Wan , Zihan Wang , Yijun Wang

Natural Language Processing (NLP) is witnessing a remarkable breakthrough driven by the success of Large Language Models (LLMs). LLMs have gained significant attention across academia and industry for their versatile applications in text…

Computation and Language · Computer Science 2024-04-16 Taojun Hu , Xiao-Hua Zhou

Large Language Models (LLMs) have demonstrated remarkable performance on various medical question-answering (QA) benchmarks, including standardized medical exams. However, correct answers alone do not ensure correct logic, and models may…

Computation and Language · Computer Science 2025-06-02 Yuexing Hao , Kumail Alhamoud , Hyewon Jeong , Haoran Zhang , Isha Puri , Philip Torr , Mike Schaekermann , Ariel D. Stern , Marzyeh Ghassemi

Large Language Models (LLMs) are increasingly adopted for applications in healthcare, reaching the performance of domain experts on tasks such as question answering and document summarisation. Despite their success on these tasks, it is…

Computation and Language · Computer Science 2025-05-20 Aishik Nagar , Viktor Schlegel , Thanh-Tung Nguyen , Hao Li , Yuping Wu , Kuluhan Binici , Stefan Winkler

Offline evaluation of search systems depends on test collections. These benchmarks provide the researchers with a corpus of documents, topics and relevance judgements indicating which documents are relevant for each topic. While test…

Information Retrieval · Computer Science 2025-07-23 David Otero , Javier Parapar , Álvaro Barreiro

Large-scale language models (LLMs) like ChatGPT have demonstrated impressive abilities in generating responses based on human instructions. However, their use in the medical field can be challenging due to their lack of specific, in-depth…

Computation and Language · Computer Science 2025-02-25 Yubo Wang , Xueguang Ma , Wenhu Chen

Clinical decisions are often required under incomplete information. Clinical experts must identify whether available information is sufficient for judgment, as both premature conclusion and unnecessary abstention can compromise patient…

Artificial Intelligence · Computer Science 2026-02-27 Yusuke Watanabe , Yohei Kobashi , Takeshi Kojima , Yusuke Iwasawa , Yasushi Okuno , Yutaka Matsuo

Continuing advances in Large Language Models (LLMs) in artificial intelligence offer important capacities in intuitively accessing and using medical knowledge in many contexts, including education and training as well as assessment and…

Computation and Language · Computer Science 2024-08-01 Roma Shusterman , Allison C. Waters , Shannon O`Neill , Phan Luu , Don M. Tucker

Objective Structured Clinical Examinations (OSCEs) are widely used to assess medical students' communication skills, but scoring interview-based assessments is time-consuming and potentially subject to human bias. This study explored the…

Computation and Language · Computer Science 2025-05-16 Jadon Geathers , Yann Hicke , Colleen Chan , Niroop Rajashekar , Justin Sewell , Susannah Cornes , Rene F. Kizilcec , Dennis Shung

Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, but the quality bar for medical and clinical applications is high. Today, attempts to assess models' clinical knowledge…

With the proliferation of Large Language Models (LLMs) in diverse domains, there is a particular need for unified evaluation standards in clinical medical scenarios, where models need to be examined very thoroughly. We present CliMedBench,…

Computation and Language · Computer Science 2024-10-07 Zetian Ouyang , Yishuai Qiu , Linlin Wang , Gerard de Melo , Ya Zhang , Yanfeng Wang , Liang He

Computed Tomography (CT) plays a crucial role in clinical diagnosis, but the growing demand for CT examinations has raised concerns about diagnostic errors. While Multimodal Large Language Models (MLLMs) demonstrate promising comprehension…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Sunggu Kyung , Hyungbin Park , Jinyoung Seo , Jimin Sung , Jihyun Kim , Dongyeong Kim , Wooyoung Jo , Yoojin Nam , Sangah Park , Taehee Kwon , Sang Min Lee , Namkug Kim

While pioneering deep learning methods have made great strides in analyzing electronic health record (EHR) data, they often struggle to fully capture the semantics of diverse medical codes from limited data. The integration of external…

Machine Learning · Computer Science 2024-08-26 Zhihao Yu , Yujie Jin , Yongxin Xu , Xu Chu , Yasha Wang , Junfeng Zhao

Existing benchmarks for evaluating mathematical reasoning in large language models (LLMs) rely primarily on competition problems, formal proofs, or artificially challenging questions -- failing to capture the nature of mathematics…

Artificial Intelligence · Computer Science 2025-10-21 Jie Zhang , Cezara Petrui , Kristina Nikolić , Florian Tramèr

The existing methods for evaluating the medical knowledge of Large Language Models (LLMs) are largely based on atemporal examination-style benchmarks, while in reality, medical knowledge is inherently dynamic and continuously evolves as new…

Machine Learning · Computer Science 2026-05-14 Zihan Guan , Qiao Jin , Guangzhi Xiong , Fangyuan Chen , Mengxuan Hu , Qingyu Chen , Yifan Peng , Zhiyong Lu , Anil Vullikanti

While automatic metrics drive progress in Machine Translation (MT) and Text Summarization (TS), existing metrics have been developed and validated almost exclusively for English and other high-resource languages. This narrow focus leaves…

Computation and Language · Computer Science 2025-10-09 Amir Hossein Yari , Kalmit Kulkarni , Ahmad Raza Khan , Fajri Koto

In recent years, Large Language Models (LLMs) have become increasingly more powerful in their ability to complete complex tasks. One such task in which LLMs are often employed is scoring, i.e., assigning a numerical value from a certain…

Computation and Language · Computer Science 2024-12-31 Henry J. Xie , Jinghan Zhang , Xinhao Zhang , Kunpeng Liu

Large Language Models (LLMs) have significantly advanced Machine Translation (MT), applying them to linguistically complex domains-such as Social Network Services, literature etc. In these scenarios, translations often require handling…

Computation and Language · Computer Science 2026-04-17 Yanzhi Tian , Cunxiang Wang , Zeming Liu , Heyan Huang , Wenbo Yu , Dawei Song , Jie Tang , Yuhang Guo