English
Related papers

Related papers: Decoding Safety Feedback from Diverse Raters: A Da…

200 papers

We evaluate how effectively platform-level parental controls moderate a mainstream conversational assistant used by minors. Our two-phase protocol first builds a category-balanced conversation corpus via PAIR-style iterative prompt…

Computers and Society · Computer Science 2026-02-02 Kerem Ersoz , Saleh Afroogh , David Atkinson , Junfeng Jiao

The validity of AI safety evaluations depends on models behaving consistently across controlled and deployment settings. Prior work has identified test-time contextual cues, such as hypothetical scenarios, as a source of verbalized…

Computation and Language · Computer Science 2026-05-28 Katharina Deckenbach , Haritz Puerto , Jonas Geiping , Sahar Abdelnabi

Incident monitoring can drive safety improvements in high-reliability industries and population-scale technologies, but remains underdeveloped in AI governance. Public databases catalog thousands of AI incidents, but simple incident counts…

Computers and Society · Computer Science 2026-05-08 Isaak Mengesha , Branwen Owen , Charlie Collins , Tina Wong , Simon Mylius , Peter Slattery , Sean McGregor

Unintended bias in Machine Learning can manifest as systemic differences in performance for different demographic groups, potentially compounding existing challenges to fairness in society at large. In this paper, we introduce a suite of…

Machine Learning · Computer Science 2019-05-09 Daniel Borkan , Lucas Dixon , Jeffrey Sorensen , Nithum Thain , Lucy Vasserman

We introduce a conceptual framework and provide considerations for the institutional design of AI incident reporting systems, i.e., processes for collecting information about safety- and rights-related events caused by general-purpose AI.…

Computers and Society · Computer Science 2026-04-15 Kevin Wei , Lennart Heim

A multitude of explainability methods and associated fidelity performance metrics have been proposed to help better understand how modern AI systems make decisions. However, much of the current work has remained theoretical -- without much…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Julien Colin , Thomas Fel , Remi Cadene , Thomas Serre

Many safety failures in machine learning arise when models are used to assign predictions to people (often in settings like lending, hiring, or content moderation) without accounting for how individuals can change their inputs. In this…

Machine Learning · Computer Science 2025-07-04 Seung Hyun Cheon , Meredith Stewart , Bogdan Kulynych , Tsui-Wei Weng , Berk Ustun

The use of large language models like ChatGPT in code review offers promising efficiency gains but also raises concerns about correctness and safety. Existing evaluation methods for code review generation either rely on automatic…

Software Engineering · Computer Science 2025-12-18 Robert Heumüller , Frank Ortmeier

From autonomous driving to package delivery, ensuring safe yet efficient multi-agent interaction is challenging as the interaction dynamics are influenced by hard-to-model factors such as social norms and contextual cues. Understanding…

Systems and Control · Electrical Eng. & Systems 2026-03-11 Isaac Remy , David Fridovich-Keil , Karen Leung

Large language models (LLMs) are currently aligned using techniques such as reinforcement learning from human feedback (RLHF). However, these methods use scalar rewards that can only reflect user preferences on average. Pluralistic…

Computation and Language · Computer Science 2025-08-13 Jadie Adams , Brian Hu , Emily Veenhuis , David Joy , Bharadwaj Ravichandran , Aaron Bray , Anthony Hoogs , Arslan Basharat

Studying the robustness of Large Language Models (LLMs) to unsafe behaviors is an important topic of research today. Building safety classification models or guard models, which are fine-tuned models for input/output safety classification…

Computation and Language · Computer Science 2025-07-30 Sowmya Vajjala

In measurement theory, instruments do not simply record reality; they help constitute what is observed. The same holds for generative AI evaluation: benchmarks do not just measure, they shape what models appear to be. Functionalist…

Artificial Intelligence · Computer Science 2026-04-23 Rebecca L. Johnson

Conversational AI systems can engage in unsafe behaviour when handling users' medical queries that can have severe consequences and could even lead to deaths. Systems therefore need to be capable of both recognising the seriousness of…

Computation and Language · Computer Science 2022-10-04 Gavin Abercrombie , Verena Rieser

In the era of widespread public use of AI systems across various domains, ensuring adversarial robustness has become increasingly vital to maintain safety and prevent undesirable errors. Researchers have curated various adversarial datasets…

Machine Learning · Computer Science 2023-11-08 Yuanchen Bai , Raoyi Huang , Vijay Viswanathan , Tzu-Sheng Kuo , Tongshuang Wu

Human-annotated data plays a critical role in the fairness of AI systems, including those that deal with life-altering decisions or moderating human-created web/social media content. Conventionally, annotator disagreements are resolved…

Information Retrieval · Computer Science 2023-07-21 Tharindu Cyril Weerasooriya , Sarah Luger , Saloni Poddar , Ashiqur R. KhudaBukhsh , Christopher M. Homan

In recent years, social media has emerged as a primary channel for users to promptly share feedback and issues during disasters and emergencies, playing a key role in crisis management. While significant progress has been made in collecting…

Computation and Language · Computer Science 2025-04-18 Loris Belcastro , Cristian Cosentino , Fabrizio Marozzo , Merve Gündüz-Cüre , Sule Öztürk-Birim

Understanding AI systems' inner workings is critical for ensuring value alignment and safety. This review explores mechanistic interpretability: reverse engineering the computational mechanisms and representations learned by neural networks…

Artificial Intelligence · Computer Science 2024-08-27 Leonard Bereska , Efstratios Gavves

AI-powered answer engines are inherently non-deterministic: identical queries submitted at different times can produce different responses and cite different sources. Despite this stochastic behavior, current approaches to measuring domain…

Applications · Statistics 2026-03-11 Ronald Sielinski

Developing robust and fair AI systems require datasets with comprehensive set of labels that can help ensure the validity and legitimacy of relevant measurements. Recent efforts, therefore, focus on collecting person-related datasets that…

Recent protocols and metrics for training and evaluating autonomous robot navigation through crowds are inconsistent due to diversified definitions of "social behavior". This makes it difficult, if not impossible, to effectively compare…

Robotics · Computer Science 2022-11-29 Junxian Wang , Wesley P. Chan , Pamela Carreno-Medrano , Akansel Cosgun , Elizabeth Croft