English
Related papers

Related papers: Securing External Deeper-than-black-box GPAI Evalu…

200 papers

The EU Artificial Intelligence (AI) Act directs businesses to assess their AI systems to ensure they are developed in a way that is human-centered and trustworthy. The rapid adoption of AI in the industry has outpaced ethical evaluation…

Computers and Society · Computer Science 2025-09-30 Louise McCormack , Diletta Huyskes , Dave Lewis , Malika Bendechache

Artificial Intelligence (AI) applications are being used to predict and assess behaviour in multiple domains, such as criminal justice and consumer finance, which directly affect human well-being. However, if AI is to improve people's…

Other Computer Science · Computer Science 2019-06-12 Andrea Aler Tubella , Andreas Theodorou , Virginia Dignum , Frank Dignum

As frontier artificial intelligence (AI) systems become more capable, it becomes more important that developers can explain why their systems are sufficiently safe. One way to do so is via safety cases: reports that make a structured…

Computers and Society · Computer Science 2024-10-30 Marie Davidsen Buhl , Gaurav Sett , Leonie Koessler , Jonas Schuett , Markus Anderljung

AI evaluations have become critical tools for assessing large language model capabilities and safety. This paper presents practical insights from eight months of maintaining $inspect\_evals$, an open-source repository of 70+…

Computation and Language · Computer Science 2025-07-10 Alexandra Abbas , Celia Waggoner , Justin Olive

The risks of frontier AI may require international cooperation, which in turn may require verification: checking that all parties follow agreed-on rules. For instance, states might need to verify that powerful AI models are widely deployed…

Computers and Society · Computer Science 2025-07-29 Mauricio Baker , Gabriel Kulp , Oliver Marks , Miles Brundage , Lennart Heim

AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210 safety benchmarks that maps out common challenges in safety…

Computers and Society · Computer Science 2026-02-10 Cheng Yu , Severin Engelmann , Ruoxuan Cao , Dalia Ali , Orestis Papakyriakopoulos

Rising concern for the societal implications of artificial intelligence systems has inspired a wave of academic and journalistic literature in which deployed systems are audited for harm by investigators from outside the organizations…

Machine learning models are becoming increasingly popular in different types of settings. This is mainly caused by their ability to achieve a level of predictive performance that is hard to match by human experts in this new era of big…

Machine Learning · Computer Science 2021-09-20 Luis Torgo , Paulo Azevedo , Ines Areosa

The rapid advancement of AI systems has raised widespread concerns about potential harms of frontier AI systems and the need for responsible evaluation and oversight. In this position paper, we argue that frontier AI companies should report…

Computers and Society · Computer Science 2025-03-25 Dillon Bowen , Ann-Kathrin Dombrowski , Adam Gleave , Chris Cundy

Interpretability of Deep Learning (DL) is a barrier to trustworthy AI. Despite great efforts made by the Explainable AI (XAI) community, explanations lack robustness -- indistinguishable input perturbations may lead to different XAI…

Machine Learning · Computer Science 2023-08-01 Wei Huang , Xingyu Zhao , Gaojie Jin , Xiaowei Huang

The rapid proliferation and deployment of General-Purpose AI (GPAI) models, including large language models (LLMs), present unprecedented challenges for AI supervisory entities. We hypothesize that these entities will need to navigate an…

Artificial Intelligence · Computer Science 2025-06-12 Manuel Cebrian , Emilia Gomez , David Fernandez Llorca

A prediction model is most useful if it generalizes beyond the development data with external validations, but to what extent should it generalize remains unclear. In practice, prediction models are externally validated using data from very…

Machine Learning · Computer Science 2023-04-11 Yilin Ning , Victor Volovici , Marcus Eng Hock Ong , Benjamin Alan Goldstein , Nan Liu

As the manufacturing industry advances with sensor integration and automation, the opaque nature of deep learning models in machine learning poses a significant challenge for fault detection and diagnosis. And despite the related predictive…

Artificial Intelligence · Computer Science 2024-06-11 Ahmed Maged , Salah Haridy , Herman Shen

The rapid advancement of ML models in critical sectors such as healthcare, finance, and security has intensified the need for robust data security, model integrity, and reliable outputs. Large multimodal foundational models, while crucial…

Cryptography and Security · Computer Science 2024-12-13 Hongyang Zhang , Yue Zhao , Claudio Angione , Harry Yang , James Buban , Ahmad Farhan , Fielding Johnston , Patrick Colangelo

XAI refers to the techniques and methods for building AI applications which assist end users to interpret output and predictions of AI models. Black box AI applications in high-stakes decision-making situations, such as medical domain have…

The expanding role of Artificial Intelligence (AI) in diverse engineering domains highlights the challenges associated with deploying AI models in new operational environments, involving substantial investments in data collection and model…

Artificial Intelligence · Computer Science 2024-05-14 Daryl Mupupuni , Anupama Guntu , Liang Hong , Kamrul Hasan , Leehyun Keel

The problem of human trust in artificial intelligence is one of the most fundamental problems in applied machine learning. Our processes for evaluating AI trustworthiness have substantial ramifications for ML's impact on science, health,…

Machine Learning · Computer Science 2022-02-14 Max W. Shen

Trust in clinical artificial intelligence (AI) cannot be reduced to model accuracy, fluency of generation, or overall positive user impression. In medicine, trust must be engineered as a measurable system property grounded in evidence,…

Computation and Language · Computer Science 2026-04-30 Serhii Zabolotnii , Viktoriia Holinko , Olha Antonenko

The black-box nature of artificial intelligence (AI) models has been the source of many concerns in their use for critical applications. Explainable Artificial Intelligence (XAI) is a rapidly growing research field that aims to create…

Cryptography and Security · Computer Science 2023-06-13 Gaith Rjoub , Jamal Bentahar , Omar Abdel Wahab , Rabeb Mizouni , Alyssa Song , Robin Cohen , Hadi Otrok , Azzam Mourad

Medical Large Language Models (LLMs) are increasingly deployed for clinical decision support across diverse specialties, yet systematic evaluation of their robustness to adversarial misuse and privacy leakage remains inaccessible to most…

Cryptography and Security · Computer Science 2025-12-10 Jinghao Wang , Ping Zhang , Carter Yagemann