English
Related papers

Related papers: Auditing Games for Sandbagging

200 papers

This study investigates the self-rationalization framework constructed with a cooperative game, where a generator initially extracts the most informative segment from raw input, and a subsequent predictor utilizes the selected subset for…

Artificial Intelligence · Computer Science 2025-08-07 Wei Liu , Zhongyu Niu , Lang Gao , Zhiying Deng , Jun Wang , Haozhao Wang , Ruixuan Li

Sustainability and efficiency have become essential considerations in the development and deployment of Artificial Intelligence systems, but existing regulatory practices for Green AI still lack standardized, model-agnostic evaluation…

Machine Learning · Computer Science 2026-03-19 Jorge Paz-Ruza , João Gama , Amparo Alonso-Betanzos , Bertha Guijarro-Berdiñas

Large Language Models (LLMs) have demonstrated strong generative capabilities but remain prone to inconsistencies and hallucinations. We introduce Peer Elicitation Games (PEG), a training-free, game-theoretic framework for aligning LLMs…

Machine Learning · Computer Science 2025-10-21 Baiting Chen , Tong Zhu , Jiale Han , Lexin Li , Gang Li , Xiaowu Dai

Customers of machine learning systems demand accountability from the companies employing these algorithms for various prediction tasks. Accountability requires understanding of system limit and condition of erroneous predictions, as…

Machine Learning · Computer Science 2021-05-12 Amita Misra , Zhe Liu , Jalal Mahmud

The rapid evolution of artificial intelligence (AI) systems, tools, and technologies has opened up novel, unprecedented opportunities for businesses to innovate, differentiate, and compete. However, growing concerns have emerged about the…

Human-Computer Interaction · Computer Science 2026-01-13 Nelly Elsayed

This study examines vulnerabilities in transformer-based automated short-answer grading systems used in medical education, with a focus on how these systems can be manipulated through adversarial gaming strategies. Our research identifies…

Computation and Language · Computer Science 2025-05-02 Sahar Yarmohammadtoosky , Yiyun Zhou , Victoria Yaneva , Peter Baldwin , Saed Rezayi , Brian Clauser , Polina Harikeo

Language model evaluations often fail to characterize consequential failure modes, forcing experts to inspect outputs and build new benchmarks. We introduce task elicitation, a method that automatically builds new evaluations to profile…

Computation and Language · Computer Science 2025-09-29 Davis Brown , Prithvi Balehannina , Helen Jin , Shreya Havaldar , Hamed Hassani , Eric Wong

Artificial Knowledge (AK) systems are transforming decision-making across critical domains such as healthcare, finance, and criminal justice. However, their growing opacity presents governance challenges that current regulatory approaches,…

Computers and Society · Computer Science 2025-05-29 Dalit Ken-Dror Feldman , Daniel Benoliel

Despite their growing use in academic writing and statistical analysis, the performance of artificial intelligence (AI) tools in scientific peer review remains a largely unexplored area. A key challenge is jagged AI, a phenomenon where AI…

Applications · Statistics 2026-05-19 Jin Wook Lee , William Szegda , Zhisheng Song , Edward L. Ionides

Red-teaming is a core part of the infrastructure that ensures that AI models do not produce harmful content. Unlike past technologies, the black box nature of generative AI systems necessitates a uniquely interactional mode of testing, one…

Despite legal mandates for the right to be forgotten, AI operators routinely fail to comply with data deletion requests. While machine unlearning (MU) provides a technical solution to remove personal data's influence from trained models,…

Machine Learning · Computer Science 2026-02-17 Qinqi Lin , Ningning Ding , Lingjie Duan , Jianwei Huang

Building reliable deception detectors for AI systems -- methods that could predict when an AI system is being strategically deceptive without necessarily requiring behavioural evidence -- would be valuable in mitigating risks from advanced…

Machine Learning · Computer Science 2025-12-17 Lewis Smith , Bilal Chughtai , Neel Nanda

In this study, we take a departure and explore an explainability-driven strategy to data auditing, where actionable insights into the data at hand are discovered through the eyes of quantitative explainability on the behaviour of a dummy…

Computer Vision and Pattern Recognition · Computer Science 2021-06-18 Alexander Wong , Adam Dorfman , Paul McInnis , Hayden Gunraj

Although deep learning has demonstrated astonishing performance in many applications, there are still concerns about its dependability. One desirable property of deep learning applications with societal impact is fairness (i.e.,…

Machine Learning · Computer Science 2021-07-30 Peixin Zhang , Jingyi Wang , Jun Sun , Xinyu Wang , Guoliang Dong , Xingen Wang , Ting Dai , Jin Song Dong

Retrieval-Augmented Generation (RAG) systems augment large language models with external knowledge, yet introduce a critical security vulnerability: RAG Knowledge Base Leakage, wherein adversarial prompts can induce the model to divulge…

Cryptography and Security · Computer Science 2026-04-14 Yuanbo Xie , Yingjie Zhang , Yulin Li , Shouyou Song , Xiaokun Chen , Zhihan Liu , Liya Su , Tingwen Liu

As machine learning systems move from computer-science laboratories into the open world, their accountability becomes a high priority problem. Accountability requires deep understanding of system behavior and its failures. Current…

Machine Learning · Computer Science 2018-09-21 Besmira Nushi , Ece Kamar , Eric Horvitz

We use an evolutionary game model to study the interplay between corporate environmental compliance and enforcement promoted by the policy maker in a country facing a pollution trap, i.e., a scenario in which the vast majority of firms do…

Physics and Society · Physics 2018-02-27 Gabriel Meyer Salomão , André Barreira da Silva Rocha

Adversarial examples can be useful for identifying vulnerabilities in AI systems before they are deployed. In reinforcement learning (RL), adversarial policies can be developed by training an adversarial agent to minimize a target agent's…

Artificial Intelligence · Computer Science 2023-10-17 Stephen Casper , Taylor Killian , Gabriel Kreiman , Dylan Hadfield-Menell

Machine learning (ML) models are known to be vulnerable to adversarial examples. Applications of ML to voice biometrics authentication are no exception. Yet, the implications of audio adversarial examples on these real-world systems remain…

As large language models are integrated into society, robustness toward a suite of prompts is increasingly important to maintain reliability in a high-variance environment.Robustness evaluations must comprehensively encapsulate the various…

Computation and Language · Computer Science 2023-11-14 Alex Mei , Sharon Levy , William Yang Wang
‹ Prev 1 8 9 10 Next ›