English
Related papers

Related papers: Auditing Games for Sandbagging

200 papers

Predictive models often reinforce biases which were originally embedded in their training data, through skewed decisions. In such cases, mitigation methods are critical to ensure that, regardless of the prevailing disparities, model…

Machine Learning · Statistics 2025-07-15 Ricardo Inácio , Zafeiris Kokkinogenis , Vitor Cerqueira , Carlos Soares

This work investigates the potential of undermining both fairness and detection performance in abusive language detection. In a dynamic and complex digital world, it is crucial to investigate the vulnerabilities of these detection models to…

Computation and Language · Computer Science 2023-12-07 Yueqing Liang , Lu Cheng , Ali Payani , Kai Shu

As AI systems become more capable and widely deployed as agents, ensuring their safe operation becomes critical. AI control offers one approach to mitigating the risk from untrusted AI agents by monitoring their actions and intervening or…

Artificial Intelligence · Computer Science 2025-11-06 Jon Kutasov , Chloe Loughridge , Yuqi Sun , Henry Sleight , Buck Shlegeris , Tyler Tracy , Joe Benton

The increasing scale and sophistication of cyberattacks has led to the adoption of machine learning based classification techniques, at the core of cybersecurity systems. These techniques promise scale and accuracy, which traditional rule…

Machine Learning · Computer Science 2018-03-28 Tegjyot Singh Sethi , Mehmed Kantardzic , Joung Woo Ryu

Botnet detection based on machine learning have witnessed significant leaps in recent years, with the availability of large and reliable datasets that are extracted from real-life scenarios. Consequently, adversarial attacks on machine…

Cryptography and Security · Computer Science 2023-10-03 Mohammed M. Alani , Atefeh Mashatan , Ali Miri

Many ML models are opaque to humans, producing decisions too complex for humans to easily understand. In response, explainable artificial intelligence (XAI) tools that analyze the inner workings of a model have been created. Despite these…

Computers and Society · Computer Science 2021-06-17 Kiana Alikhademi , Brianna Richardson , Emma Drobina , Juan E. Gilbert

White-box monitors are a popular technique for detecting potentially harmful behaviours in language models. While they perform well in general, their effectiveness in detecting text-ambiguous behaviour is disputed. In this work, we find…

Artificial Intelligence · Computer Science 2026-03-10 Gerard Boxo , Aman Neelappa , Shivam Raval

While there have been many proposals on making AI algorithms explainable, few have attempted to evaluate the impact of AI-generated explanations on human performance in conducting human-AI collaborative tasks. To bridge the gap, we propose…

Computers and Society · Computer Science 2019-09-24 Arijit Ray , Yi Yao , Rakesh Kumar , Ajay Divakaran , Giedrius Burachas

Responsible AI is becoming critical as AI is widely used in our everyday lives. Many companies that deploy AI publicly state that when training a model, we not only need to improve its accuracy, but also need to guarantee that the model…

Machine Learning · Computer Science 2021-01-18 Steven Euijong Whang , Ki Hyun Tae , Yuji Roh , Geon Heo

Edge storage presents a viable data storage alternative for application vendors (AV), offering benefits such as reduced bandwidth overhead and latency compared to cloud storage. However, data cached in edge computing systems is susceptible…

Cryptography and Security · Computer Science 2023-12-27 Zahra Seyedi , Farhad Rahmati , Mohammad Ali , Ximeng Liu

To safely deploy deep learning-based computer vision models for computer-aided detection and diagnosis, we must ensure that they are robust and reliable. Towards that goal, algorithmic auditing has received substantial attention. To guide…

Machine Learning · Computer Science 2023-04-07 Mitchell Pavlak , Nathan Drenkow , Nicholas Petrick , Mohammad Mehdi Farhangi , Mathias Unberath

Agent skills extend LLM agents with reusable instructions, tool interfaces, and executable code, and users increasingly install third-party skills from marketplaces, repositories, and community channels. Because a skill exposes both…

Cryptography and Security · Computer Science 2026-05-13 Zhaojiacheng Zhou

A concern about cutting-edge or "frontier" AI foundation models is that an adversary may use the models for preparing chemical, biological, radiological, nuclear, (CBRN), cyber, or other attacks. At least two methods can identify foundation…

Cryptography and Security · Computer Science 2024-05-21 Anthony M. Barrett , Krystal Jackson , Evan R. Murphy , Nada Madkour , Jessica Newman

Artificial intelligence-based systems for player risk detection have become central to harm prevention efforts in the gambling industry. However, growing concerns around transparency and effectiveness have highlighted the absence of…

Debugging is a crucial skill in programming education and software development, yet it is often overlooked in CS curricula. To address this, we introduce an AI-powered debugging assistant integrated into an IDE. It offers real-time support…

The deployment of AI systems in safety-critical domains, such as industrial defect inspection, autonomous driving, and medical diagnosis, is severely hampered by their lack of reliability. A single undetected erroneous prediction can lead…

Computer Vision and Pattern Recognition · Computer Science 2026-04-22 Hang-Cheng Dong , Yuhao Jiang , Yibo Jiao , Lu Zou , Kai Zheng , Bingguo Liu , Dong Ye , Guodong Liu

Recent AI-related scandals have shed a spotlight on accountability in AI, with increasing public interest and concern. This paper draws on literature from public policy and governance to make two contributions. First, we propose an AI…

Computers and Society · Computer Science 2021-10-19 Chris Percy , Simo Dragicevic , Sanjoy Sarkar , Artur S. d'Avila Garcez

It is important to collect credible training samples $(x,y)$ for building data-intensive learning systems (e.g., a deep learning system). Asking people to report complex distribution $p(x)$, though theoretically viable, is challenging in…

Machine Learning · Computer Science 2021-02-26 Jiaheng Wei , Zuyue Fu , Yang Liu , Xingyu Li , Zhuoran Yang , Zhaoran Wang

Fairness in artificial intelligence (AI) prediction models is increasingly emphasized to support responsible adoption in high-stakes domains such as health care and criminal justice. Guidelines and implementation frameworks highlight the…

Machine Learning · Computer Science 2025-04-14 Yilin Ning , Yian Ma , Mingxuan Liu , Xin Li , Nan Liu

Benchmarks are important measures to evaluate safety and compliance of AI models at scale. However, they typically do not offer verifiable results and lack confidentiality for model IP and benchmark datasets. We propose Attestable Audits,…

Artificial Intelligence · Computer Science 2025-07-01 Christoph Schnabl , Daniel Hugenroth , Bill Marino , Alastair R. Beresford