中文
相关论文

相关论文: Securing External Deeper-than-black-box GPAI Evalu…

200 篇论文

The EU Artificial Intelligence (AI) Act directs businesses to assess their AI systems to ensure they are developed in a way that is human-centered and trustworthy. The rapid adoption of AI in the industry has outpaced ethical evaluation…

计算机与社会 · 计算机科学 2025-09-30 Louise McCormack , Diletta Huyskes , Dave Lewis , Malika Bendechache

Artificial Intelligence (AI) applications are being used to predict and assess behaviour in multiple domains, such as criminal justice and consumer finance, which directly affect human well-being. However, if AI is to improve people's…

其他计算机科学 · 计算机科学 2019-06-12 Andrea Aler Tubella , Andreas Theodorou , Virginia Dignum , Frank Dignum

As frontier artificial intelligence (AI) systems become more capable, it becomes more important that developers can explain why their systems are sufficiently safe. One way to do so is via safety cases: reports that make a structured…

计算机与社会 · 计算机科学 2024-10-30 Marie Davidsen Buhl , Gaurav Sett , Leonie Koessler , Jonas Schuett , Markus Anderljung

AI evaluations have become critical tools for assessing large language model capabilities and safety. This paper presents practical insights from eight months of maintaining $inspect\_evals$, an open-source repository of 70+…

计算与语言 · 计算机科学 2025-07-10 Alexandra Abbas , Celia Waggoner , Justin Olive

The risks of frontier AI may require international cooperation, which in turn may require verification: checking that all parties follow agreed-on rules. For instance, states might need to verify that powerful AI models are widely deployed…

计算机与社会 · 计算机科学 2025-07-29 Mauricio Baker , Gabriel Kulp , Oliver Marks , Miles Brundage , Lennart Heim

AI safety benchmarks are pivotal for safety in advanced AI systems; however, they have significant technical, epistemic, and sociotechnical shortcomings. We present a review of 210 safety benchmarks that maps out common challenges in safety…

计算机与社会 · 计算机科学 2026-02-10 Cheng Yu , Severin Engelmann , Ruoxuan Cao , Dalia Ali , Orestis Papakyriakopoulos

Rising concern for the societal implications of artificial intelligence systems has inspired a wave of academic and journalistic literature in which deployed systems are audited for harm by investigators from outside the organizations…

Machine learning models are becoming increasingly popular in different types of settings. This is mainly caused by their ability to achieve a level of predictive performance that is hard to match by human experts in this new era of big…

机器学习 · 计算机科学 2021-09-20 Luis Torgo , Paulo Azevedo , Ines Areosa

The rapid advancement of AI systems has raised widespread concerns about potential harms of frontier AI systems and the need for responsible evaluation and oversight. In this position paper, we argue that frontier AI companies should report…

计算机与社会 · 计算机科学 2025-03-25 Dillon Bowen , Ann-Kathrin Dombrowski , Adam Gleave , Chris Cundy

Interpretability of Deep Learning (DL) is a barrier to trustworthy AI. Despite great efforts made by the Explainable AI (XAI) community, explanations lack robustness -- indistinguishable input perturbations may lead to different XAI…

机器学习 · 计算机科学 2023-08-01 Wei Huang , Xingyu Zhao , Gaojie Jin , Xiaowei Huang

The rapid proliferation and deployment of General-Purpose AI (GPAI) models, including large language models (LLMs), present unprecedented challenges for AI supervisory entities. We hypothesize that these entities will need to navigate an…

人工智能 · 计算机科学 2025-06-12 Manuel Cebrian , Emilia Gomez , David Fernandez Llorca

A prediction model is most useful if it generalizes beyond the development data with external validations, but to what extent should it generalize remains unclear. In practice, prediction models are externally validated using data from very…

机器学习 · 计算机科学 2023-04-11 Yilin Ning , Victor Volovici , Marcus Eng Hock Ong , Benjamin Alan Goldstein , Nan Liu

As the manufacturing industry advances with sensor integration and automation, the opaque nature of deep learning models in machine learning poses a significant challenge for fault detection and diagnosis. And despite the related predictive…

人工智能 · 计算机科学 2024-06-11 Ahmed Maged , Salah Haridy , Herman Shen

The rapid advancement of ML models in critical sectors such as healthcare, finance, and security has intensified the need for robust data security, model integrity, and reliable outputs. Large multimodal foundational models, while crucial…

密码学与安全 · 计算机科学 2024-12-13 Hongyang Zhang , Yue Zhao , Claudio Angione , Harry Yang , James Buban , Ahmad Farhan , Fielding Johnston , Patrick Colangelo

XAI refers to the techniques and methods for building AI applications which assist end users to interpret output and predictions of AI models. Black box AI applications in high-stakes decision-making situations, such as medical domain have…

The expanding role of Artificial Intelligence (AI) in diverse engineering domains highlights the challenges associated with deploying AI models in new operational environments, involving substantial investments in data collection and model…

人工智能 · 计算机科学 2024-05-14 Daryl Mupupuni , Anupama Guntu , Liang Hong , Kamrul Hasan , Leehyun Keel

The problem of human trust in artificial intelligence is one of the most fundamental problems in applied machine learning. Our processes for evaluating AI trustworthiness have substantial ramifications for ML's impact on science, health,…

机器学习 · 计算机科学 2022-02-14 Max W. Shen

Trust in clinical artificial intelligence (AI) cannot be reduced to model accuracy, fluency of generation, or overall positive user impression. In medicine, trust must be engineered as a measurable system property grounded in evidence,…

计算与语言 · 计算机科学 2026-04-30 Serhii Zabolotnii , Viktoriia Holinko , Olha Antonenko

The black-box nature of artificial intelligence (AI) models has been the source of many concerns in their use for critical applications. Explainable Artificial Intelligence (XAI) is a rapidly growing research field that aims to create…

密码学与安全 · 计算机科学 2023-06-13 Gaith Rjoub , Jamal Bentahar , Omar Abdel Wahab , Rabeb Mizouni , Alyssa Song , Robin Cohen , Hadi Otrok , Azzam Mourad

Medical Large Language Models (LLMs) are increasingly deployed for clinical decision support across diverse specialties, yet systematic evaluation of their robustness to adversarial misuse and privacy leakage remains inaccessible to most…

密码学与安全 · 计算机科学 2025-12-10 Jinghao Wang , Ping Zhang , Carter Yagemann