English
Related papers

Related papers: GAPS: A Clinically Grounded, Automated Benchmark f…

200 papers

In high-stakes domains like legal question-answering, the accuracy and trustworthiness of generative AI systems are of paramount importance. This work presents a comprehensive benchmark of various methods to assess the groundedness of…

Computation and Language · Computer Science 2024-10-14 Dietrich Trautmann , Natalia Ostapuk , Quentin Grail , Adrian Alan Pol , Guglielmo Bonifazi , Shang Gao , Martin Gajek

The rapid advancement of General Purpose AI (GPAI) models necessitates robust evaluation frameworks, especially with emerging regulations like the EU AI Act and its associated Code of Practice (CoP). Current AI evaluation practices depend…

Artificial Intelligence · Computer Science 2025-08-11 Matteo Prandi , Vincenzo Suriani , Federico Pierucci , Marcello Galisai , Daniele Nardi , Piercosma Bisconti

The recent years witness a trend of applying large-scale distributed deep learning in both business and scientific computing areas, whose goal is to speed up the training time to achieve a state-of-the-art quality. The HPC community feels a…

Performance · Computer Science 2020-07-02 Zihan Jiang , Lei Wang , Xingwang Xiong , Wanling Gao , Chunjie Luo , Fei Tang , Chuanxin Lan , Hongxiao Li , Jianfeng Zhan

Aggregate accuracy metrics dominate the evaluation of clinical AI decision-support systems but do not detect deployment-phase failures of input reliability, subgroup equity, threshold sensitivity, or operational feasibility. We propose the…

Machine Learning · Computer Science 2026-05-14 Rohith Reddy Bellibatlu

HealthBench, a benchmark designed to measure the capabilities of AI systems for health better (Arora et al., 2025), has advanced medical language model evaluation through physician-crafted dialogues and transparent rubrics. However, its…

Artificial Intelligence · Computer Science 2025-08-04 Fred Mutisya , Shikoh Gitau , Nasubo Ongoma , Keith Mbae , Elizabeth Wamicha

Despite the rapid progress of Multimodal Large Language Models (MLLMs), their ability to perform reliable visual grounding in high-stakes clinical software environments remains underexplored. Existing GUI benchmarks largely focus on…

Computer Vision and Pattern Recognition · Computer Science 2026-03-23 Rozain Shakeel , Abdul Rahman Mohammad Ali , Muneeb Mushtaq , Tausifa Jan Saleem , Tajamul Ashraf

AI models are increasingly deployed in live clinical environments where they must perform reliably across complex, high-stakes workflows that standard training and validation datasets were never designed to capture. Evaluating these systems…

Artificial Intelligence · Computer Science 2026-05-12 Prasanna Desikan , Harshit Rajgarhia , Shivali Dalmia , Ananya Mantravadi

The competency of any intelligent agent is bounded by its formal account of the world in which it operates. Clinical AI lacks such an account. Existing frameworks address evaluation, regulation, or system design in isolation, without a…

Background: When selecting predictive tools, for implementation in clinical practice or for recommendation in guidelines, clinicians are challenged with an overwhelming and ever-growing number of tools. Many of these have never been…

Computers and Society · Computer Science 2019-07-29 Mohamed Khalifa , Farah Magrabi , Blanca Gallego

Advances in generative AI point towards a new era of personalized applications that perform diverse tasks on behalf of users. While general AI assistants have yet to fully emerge, their potential to share personal data raises significant…

Artificial Intelligence · Computer Science 2024-09-24 Zhao Cheng , Diane Wan , Matthew Abueg , Sahra Ghalebikesabi , Ren Yi , Eugene Bagdasarian , Borja Balle , Stefan Mellem , Shawn O'Banion

The ability to research and synthesize knowledge is central to human expertise and progress. A new class of AI systems--designed for generative research synthesis--aims to automate this process by retrieving information from the live web…

Computation and Language · Computer Science 2026-02-10 Liana Patel , Negar Arabzadeh , Harshit Gupta , Ankita Sundar , Ion Stoica , Matei Zaharia , Carlos Guestrin

While the capabilities and utility of AI systems have advanced, rigorous norms for evaluating these systems have lagged. Grand claims, such as models achieving general reasoning capabilities, are supported with model performance on narrow…

Ensuring safe and effective use of AI requires understanding and anticipating its performance on novel tasks, from advanced scientific challenges to transformed workplace activities. So far, benchmarking has guided progress in AI, but it…

As LLMs achieved breakthroughs in general reasoning, their proficiency in specialized scientific domains reveals pronounced gaps in existing benchmarks due to data contamination, insufficient complexity, and prohibitive human labor costs.…

Artificial Intelligence · Computer Science 2026-02-27 Peiyao Xiao , Xiaogang Li , Chengliang Xu , Jiayi Wang , Ben Wang , Zichao Chen , Zeyu Wang , Kejun Yu , Yueqian Chen , Xulin Liu , Wende Xiao , Bing Zhao , Hu Wei

Objective: To determine the completeness of argumentative steps necessary to conclude effectiveness of an algorithm in a sample of current ML/AI supervised learning literature. Data Sources: Papers published in the Neural Information…

Machine Learning · Computer Science 2018-12-19 Franz J Király , Bilal Mateen , Raphael Sonabend

Human-in-the-loop validation is essential in safety-critical clinical AI, yet the transition between initial model inference and expert correction is rarely analyzed as a structured signal. We introduce a diagnostic alignment framework in…

Artificial Intelligence · Computer Science 2026-02-27 Dimitrios P. Panagoulias , Evangelia-Aikaterini Tsichrintzi , Georgios Savvidis , Evridiki Tsoureli-Nikita

Background: Clinical trials rely on transparent inclusion criteria to ensure generalizability. In contrast, benchmarks validating health-related large language models (LLMs) rarely characterize the "patient" or "query" populations they…

Artificial Intelligence · Computer Science 2026-04-17 Alvin Rajkomar , Pavan Sudarshan , Angela Lai , Lily Peng

Objective. Clinical AI documentation systems require evaluation methodologies that are clinically valid, economically viable, and sensitive to iterative changes. Methods requiring expert review per scoring instance are too slow and…

Artificial Intelligence · Computer Science 2026-04-28 Aaryan Shah , Andrew Hines , Alexia Downs , Denis Bajet , Paulius Mui , Fabiano Araujo , Laura Offutt , Aida Rutledge , Elizabeth Jimenez

Generative Artificial Intelligence (AI) holds immense potential in medical applications. Numerous studies have explored the efficacy of various generative AI models within healthcare contexts, but there is a lack of a comprehensive and…

Human-Computer Interaction · Computer Science 2023-12-19 Jinghong Chen , Lingxuan Zhu , Weiming Mou , Zaoqu Liu , Quan Cheng , Anqi Lin , Jian Zhang , Peng Luo

Understanding research papers remains challenging for foundation models due to specialized scientific discourse and complex figures and tables, yet existing benchmarks offer limited fine-grained evaluation at scale. To address this gap, we…

Computation and Language · Computer Science 2026-05-01 Yelin Chen , Fanjin Zhang , Suping Sun , Yunhe Pang , Yuanchun Wang , Jian Song , Xiaoyan Li , Lei Hou , Shu Zhao , Jie Tang , Juanzi Li
‹ Prev 1 2 3 10 Next ›