中文
相关论文

相关论文: Safetywashing: Do AI Safety Benchmarks Actually Me…

200 篇论文

The creation of benchmarks to evaluate the safety of Large Language Models is one of the key activities within the trusted AI community. These benchmarks allow models to be compared for different aspects of safety such as toxicity, bias,…

人工智能 · 计算机科学 2025-06-23 Lina Berrayana , Sean Rooney , Luis Garcés-Erice , Ioana Giurgiu

Rapidly advancing artificial intelligence (AI) systems introduce novel, uncertain, and potentially catastrophic risks. Managing these risks requires a mature risk-management infrastructure whose cornerstone is rigorous risk modeling. We…

The development of privacy-enhancing technologies has made immense progress in reducing trade-offs between privacy and performance in data exchange and analysis. Similar tools for structured transparency could be useful for AI governance by…

人工智能 · 计算机科学 2023-03-22 Emma Bluemke , Tantum Collins , Ben Garfinkel , Andrew Trask

The advent of advanced AI underscores the urgent need for comprehensive safety evaluations, necessitating collaboration across communities (i.e., AI, software engineering, and governance). However, divergent practices and terminologies…

软件工程 · 计算机科学 2024-05-17 Boming Xia , Qinghua Lu , Liming Zhu , Zhenchang Xing

The increasing integration of Artificial Intelligence across multiple industry sectors necessitates robust mechanisms for ensuring transparency, trust, and auditability of its development and deployment. This topic is particularly important…

密码学与安全 · 计算机科学 2025-03-31 Kar Balan , Robert Learney , Tim Wood

The deployment of artificial intelligence (AI) applications has accelerated rapidly. AI enabled technologies are facing the public in many ways including infrastructure, consumer products and home applications. Because many of these…

人工智能 · 计算机科学 2024-08-01 Joanna F. DeFranco , Luke Biersmith

We draw on our experience working on system and software assurance and evaluation for systems important to society to summarise how safety engineering is performed in traditional critical systems, such as aircraft flight control. We analyse…

计算机与社会 · 计算机科学 2025-02-07 Robin Bloomfield , John Rushby

The use of artificial intelligence (AI) and AI methods in the workplace holds both great opportunities as well as risks to occupational safety and discrimination. In addition to legal regulation, technical standards will play a key role in…

计算机与社会 · 计算机科学 2021-08-27 Nikolas Becker , Pauline Junginger , Lukas Martinez , Daniel Krupka , Leonie Beining

This paper presents an argument that certain AI safety measures, rather than mitigating existential risk, may instead exacerbate it. Under certain key assumptions - the inevitability of AI failure, the expected correlation between an AI…

人工智能 · 计算机科学 2024-06-04 Herman Cappelen , Josh Dever , John Hawthorne

To make deliberate progress towards more intelligent and more human-like artificial systems, we need to be following an appropriate feedback signal: we need to be able to define and evaluate intelligence in a way that enables comparisons…

人工智能 · 计算机科学 2019-11-26 François Chollet

This work examines an imbalance in artificial intelligence (AI) security research: the field tends to produce more work on attacking AI systems than on defending them. Drawing on related academic papers, we find biased attack-to-defense…

密码学与安全 · 计算机科学 2026-05-25 Youqian Zhang

Although AI systems have been applied in various fields and achieved impressive performance, their safety and reliability are still a big concern. This is especially important for safety-critical tasks. One shared characteristic of these…

人工智能 · 计算机科学 2023-08-08 Shuang Ao

Today's Internet Services are undergoing fundamental changes and shifting to an intelligent computing era where AI is widely employed to augment services. In this context, many innovative AI algorithms, systems, and architectures are…

The meteoric rise of AI, with its rapidly expanding market capitalization, presents both transformative opportunities and critical challenges. Chief among these is the urgent need for a new, unified paradigm for trustworthy evaluation, as…

AI evaluations are an important component of the AI governance toolkit, underlying current approaches to safety cases for preventing catastrophic risks. Our paper examines what these evaluations can and cannot tell us. Evaluations can…

计算机与社会 · 计算机科学 2024-12-13 Peter Barnett , Lisa Thiergart

In the past few decades, artificial intelligence (AI) technology has experienced swift developments, changing everyone's daily life and profoundly altering the course of human society. The intention of developing AI is to benefit humans, by…

人工智能 · 计算机科学 2021-08-20 Haochen Liu , Yiqi Wang , Wenqi Fan , Xiaorui Liu , Yaxin Li , Shaili Jain , Yunhao Liu , Anil K. Jain , Jiliang Tang

Artificial intelligence (AI) systems connected to sensor-laden devices are becoming pervasive, which has notable implications for a range of AI risks, including to privacy, the environment, autonomy and more. There is therefore a growing…

计算机与社会 · 计算机科学 2025-03-25 Mona Sloane , Emanuel Moss , Susan Kennedy , Matthew Stewart , Pete Warden , Brian Plancher , Vijay Janapa Reddi

Various AI safety datasets have been developed to measure LLMs against evolving interpretations of harm. Our evaluation of five recently published open-source safety benchmarks reveals distinct semantic clusters using UMAP dimensionality…

机器学习 · 计算机科学 2025-05-26 Jonathan Bennion , Shaona Ghosh , Mantek Singh , Nouha Dziri

In this paper we present the first steps towards hardening the science of measuring AI systems, by adopting metrology, the science of measurement and its application, and applying it to human (crowd) powered evaluations. We begin with the…

人工智能 · 计算机科学 2019-11-06 Chris Welty , Praveen Paritosh , Lora Aroyo

Open-weight advanced AI models -- systems whose parameters are freely available for download and adaptation -- are reshaping the global AI landscape. As these models rapidly close the performance gap with closed alternatives, they enable…

计算机与社会 · 计算机科学 2026-02-24 Bengüsu Özcan , Alex Petropoulos , Max Reddel