English
Related papers

Related papers: RISED: A Pre-Deployment Safety Evaluation Framewor…

200 papers

Background: Clinical trials rely on transparent inclusion criteria to ensure generalizability. In contrast, benchmarks validating health-related large language models (LLMs) rarely characterize the "patient" or "query" populations they…

Artificial Intelligence · Computer Science 2026-04-17 Alvin Rajkomar , Pavan Sudarshan , Angela Lai , Lily Peng

As AI becomes prevalent in high-risk domains and decision-making, it is essential to test for potential harms and biases. This urgency is reflected by the global emergence of AI regulations that emphasise fairness and adequate testing, with…

Machine Learning · Computer Science 2025-07-25 Varsha Ramineni , Hossein A. Rahmani , Emine Yilmaz , David Barber

Artificial Intelligence (AI) has demonstrated potential in healthcare, particularly in enhancing diagnostic accuracy and decision-making through Clinical Decision Support Systems (CDSSs). However, the successful implementation of these…

Human-Computer Interaction · Computer Science 2025-01-29 Olya Rezaeian , Alparslan Emrah Bayrak , Onur Asan

Medical imaging AI development is fundamentally dependent on annotated datasets, yet no existing standard provides machine-enforceable validation across dataset structure, annotation provenance, quality documentation, and ML readiness…

Image and Video Processing · Electrical Eng. & Systems 2026-04-21 Joan S. Muthu , John Shalen

With growing concerns regarding bias and discrimination in predictive models, the AI community has increasingly focused on assessing AI system trustworthiness. Conventionally, trustworthy AI literature relies on the probabilistic framework…

Machine Learning · Statistics 2024-01-05 Ritwik Vashistha , Arya Farahi

Widespread adoption of AI for medical decision making is still hindered due to ethical and safety-related concerns. For AI-based decision support systems in healthcare settings it is paramount to be reliable and trustworthy. Common deep…

Machine Learning · Computer Science 2024-01-26 Adrian Lindenmeyer , Malte Blattmann , Stefan Franke , Thomas Neumuth , Daniel Schneider

The emergence of autonomous, high-velocity Agentic AI systems is creating an internal assurance scalability crisis. Point-in-time, document-based audits cannot keep pace with non deterministic behaviour and distributed deployments of agents…

Computers and Society · Computer Science 2026-03-05 Guy Lupo , Bao Quoc Vo , Natania Locke

Modern artificial intelligence governance lacks a formal, enforceable mechanism for determining whether a given AI system is legally permitted to operate in a specific domain and jurisdiction. Existing tools such as model cards, audits, and…

Computers and Society · Computer Science 2026-01-15 Daniel Djan Saparning

Generative AI is rapidly moving from research to deployment, elevating the need for responsible development, evaluation, and governance. We conduct a PRISMA guided review of 232 studies (November 2022 - December 2025), spanning large…

Artificial intelligence (AI) offers incredible possibilities for patient care, but raises significant ethical issues, such as the potential for bias. Powerful ethical frameworks exist to minimize these issues, but are often developed for…

Computers and Society · Computer Science 2025-07-04 Ion Nemteanu , Adir Mancebo , Leslie Joe , Ryan Lopez , Patricia Lopez , Warren Woodrich Pettine

The rapid integration of AI into education has prioritized capability over trustworthiness, creating significant risks. Real-world deployments reveal that even advanced models are insufficient without extensive architectural scaffolding to…

Computers and Society · Computer Science 2026-01-13 Abu Syed

Artificial intelligence in high-stakes tabular domains cannot be evaluated by predictive performance alone, yet current practice still assesses explainability, fairness, robustness, privacy, and sustainability mostly in isolation. We…

Machine Learning · Computer Science 2026-05-15 Phuc Truong Loc Nguyen , Thanh Hung Do , Truong Thanh Hung Nguyen , Hung Cao

The recent success of machine learning methods applied to time series collected from Intensive Care Units (ICU) exposes the lack of standardized machine learning benchmarks for developing and comparing such methods. While raw datasets, such…

Machine Learning · Computer Science 2022-01-19 Hugo Yèche , Rita Kuznetsova , Marc Zimmermann , Matthias Hüser , Xinrui Lyu , Martin Faltys , Gunnar Rätsch

As AI models scale to billions of parameters and operate with increasing autonomy, ensuring their safe, reliable operation demands engineering-grade security and assurance frameworks. This paper presents an enterprise-level, risk-aware,…

Cryptography and Security · Computer Science 2025-05-13 Krti Tallam

The integration of AI into radiology introduces opportunities for improved clinical care provision and efficiency but it demands a meticulous approach to mitigate potential risks as with any other new technology. Beginning with rigorous…

This paper presents a dataset, called Reeds, for research on robot perception algorithms. The dataset aims to provide demanding benchmark opportunities for algorithms, rather than providing an environment for testing application-specific…

Computer Vision and Pattern Recognition · Computer Science 2021-09-20 Ola Benderius , Christian Berger , Krister Blanch

Automated content analysis increasingly supports communication research, yet scaling manual coding into computational pipelines raises concerns about measurement reliability and validity. We introduce a Hierarchical Error Correction (HEC)…

Computation and Language · Computer Science 2025-10-27 Zhilong Zhao , Yindi Liu

Performance uncertainty quantification is essential for reliable validation and eventual clinical translation of medical imaging artificial intelligence (AI). Confidence intervals (CIs) play a central role in this process by indicating how…

Deep learning models are increasingly used in scientific prediction tasks where strong benchmark performance is often interpreted as evidence of scientifically meaningful behavior. This interpretation is fragile, as models may exploit…

Machine Learning · Computer Science 2026-05-22 Barbara Tarantino , Gennaro Auricchio , Paolo Giudici

Progress in Prognostics and Health Management (PHM) is hindered by the lack of standardized and reusable evaluation practices across tasks, datasets, and application domains. Reported results are often difficult to reproduce and compare, as…

Artificial Intelligence · Computer Science 2026-05-28 Lev Telyatnikov , Raffael Theiler , Leandro Von Krannichfeldt , Olga Fink