English
Related papers

Related papers: GAPS: A Clinically Grounded, Automated Benchmark f…

200 papers

Recent years witness a trend of applying large-scale distributed deep learning algorithms (HPC AI) in both business and scientific computing areas, whose goal is to speed up the training time to achieve a state-of-the-art quality. The HPC…

Performance · Computer Science 2021-02-26 Zihan Jiang , Wanling Gao , Fei Tang , Xingwang Xiong , Lei Wang , Chuanxin Lan , Chunjie Luo , Hongxiao Li , Jianfeng Zhan

As AI-driven document understanding and processing tools become increasingly prevalent in real-world applications, the need for rigorous evaluation standards has grown increasingly urgent. Existing benchmarks and evaluations often focus on…

The reliability of clinical artificial intelligence (AI) depends on high-quality data, yet Electronic Health Records are often inconsistent with existing scientific knowledge. Current quality assessments are limited: they either focus on…

Machine Learning · Computer Science 2026-05-26 Irena Girshovitz , Atai Ambus , Moni Shahar , Ran Gilad-Bachrach

As autonomous AI agents are increasingly deployed in high-stakes environments, ensuring their safety and alignment with human values is becoming a practical deployment concern. Current benchmarks for AI agents primarily evaluate refusal of…

Artificial Intelligence · Computer Science 2026-05-12 Miles Q. Li , Benjamin C. M. Fung , Martin Weiss , Pulei Xiong , Khalil Al-Hussaeni , Claude Fachkha

Background and Objectives: Clinical Practice Guidelines (CPGs) represent the foremost methodology for sharing state-of-the-art research findings in the healthcare domain with medical practitioners to limit practice variations, reduce…

Artificial Intelligence · Computer Science 2020-12-11 Musarrat Hussain , Jamil Hussain , Taqdir Ali , Fahad Ahmed Satti , Sungyoung Lee

This paper introduces BioAgent Bench, a benchmark dataset and an evaluation suite designed for measuring the performance and robustness of AI agents in common bioinformatics tasks. The benchmark contains curated end-to-end tasks (e.g.,…

Artificial Intelligence · Computer Science 2026-05-08 Dionizije Fa , Marko Culjak , Bruno Pandza , Mateo Cupic

Clinical reasoning in medicine is a hypothesis-driven process where physicians refine diagnoses from limited information through targeted history, physical examination, and diagnostic investigations. In contrast, current medical benchmarks…

Machine Learning · Computer Science 2025-10-14 Christopher Chiu , Silviu Pitis , Mihaela van der Schaar

As dialogue systems and chatbots increasingly integrate into everyday interactions, the need for efficient and accurate evaluation methods becomes paramount. This study explores the comparative performance of human and AI assessments across…

Computation and Language · Computer Science 2024-09-11 Ike Ebubechukwu , Johane Takeuchi , Antonello Ceravola , Frank Joublin

Generative artificial intelligence (AI) offers scalable support for formative feedback, yet most AI-generated feedback relies on task-specific rubrics authored by domain experts. While effective, rubric authoring is time-consuming and…

Computation and Language · Computer Science 2026-04-15 Xin Xia , Nejla Yuruk , Yun Wang , Xiaoming Zhai

We introduce GUI-360$^\circ$, a large-scale, comprehensive dataset and benchmark suite designed to advance computer-using agents (CUAs). CUAs present unique challenges and is constrained by three persistent gaps: a scarcity of real-world…

Ensuring web accessibility is crucial for advancing social welfare, justice, and equality in digital spaces, yet the vast majority of website user interfaces remain non-compliant, due in part to the resource-intensive and unscalable nature…

Artificial Intelligence · Computer Science 2025-11-06 Ming Gu , Ziwei Wang , Sicen Lai , Zirui Gao , Sheng Zhou , Jiajun Bu

Reasoning benchmarks measure clinical performance on clean inputs. We evaluate the step before reasoning: retrieval over real EHR notes, where negation, temporality, and family-versus-patient attribution can flip a correct answer to a wrong…

Computation and Language · Computer Science 2026-05-13 Alex Stinard

The past years have presented a surge in (AI) development, fueled by breakthroughs in deep learning, increased computational power, and substantial investments in the field. Given the generative capabilities of more recent AI systems, the…

Existing automated research systems operate as stateless, linear pipelines -- generating outputs without maintaining any persistent understanding of the research landscape they navigate. They process papers sequentially, propose ideas…

Artificial Intelligence · Computer Science 2026-03-27 Yunbo Long

Rigorous evaluation of domain-specific language models requires benchmarks that are comprehensive, contamination-resistant, and maintainable. Static, manually curated datasets do not satisfy these properties. We present a graph-based…

Artificial Intelligence · Computer Science 2026-05-18 Jessica M. Lundin , Usman Nasir Nakakana , Guillaume Chabot-Couture

This document presents a preliminary compilation of general-purpose AI (GPAI) evaluation practices that may promote internal validity, external validity and reproducibility. It includes suggestions for human uplift studies and benchmark…

Computers and Society · Computer Science 2025-08-20 Patricia Paskov , Michael J. Byun , Kevin Wei , Toby Webster

Background. When selecting predictive tools, clinicians and healthcare professionals are challenged with an overwhelming number of tools, most of which have never been evaluated for comparative effectiveness. To overcome this challenge, the…

Computers and Society · Computer Science 2019-07-29 Mohamed Khalifa , Farah Magrabi , Blanca Gallego

The rapid proliferation of Generative AI (GenAI) into diverse, high-stakes domains necessitates robust and reproducible evaluation methods. However, practitioners often resort to ad-hoc, non-standardized scripts, as common metrics are often…

Computation and Language · Computer Science 2026-03-24 Nitin Gupta , Pallav Koppisetti , Kausik Lakkaraju , Biplav Srivastava

Drug development and pharmacovigilance are frequently bottlenecked by legacy clinical reporting pipelines. These monolithic systems encode regulatory-grade logic but resist AI integration by producing opaque output with no machine-readable…

Software Engineering · Computer Science 2026-05-15 Jaime Yan

The integration of artificial intelligence (AI) into medical diagnostic workflows requires robust and consistent evaluation methods to ensure reliability, clinical relevance, and the inherent variability in expert judgments. Traditional…