English
Related papers

Related papers: Decoding Safety Feedback from Diverse Raters: A Da…

200 papers

Addressing public safety effectively requires incorporating diverse stakeholder perspectives, particularly those of the community, which are often underrepresented compared to other stakeholders. This study presents a comprehensive analysis…

We evaluate the robustness of several large language models on multiple datasets. Robustness here refers to the relative insensitivity of the model's answers to meaning-preserving variants of their input. Benchmark datasets are constructed…

Computation and Language · Computer Science 2024-11-05 Samuel Ackerman , Ella Rabinovich , Eitan Farchi , Ateret Anaby-Tavor

Incorporating every annotator's perspective is crucial for unbiased data modeling. Annotator fatigue and changing opinions over time can distort dataset annotations. To combat this, we propose to learn a more accurate representation of…

Machine Learning · Computer Science 2024-06-05 Uthman Jinadu , Yi Ding

With the rapid evolution of large language models (LLMs), new and hard-to-predict harmful capabilities are emerging. This requires developers to be able to identify risks through the evaluation of "dangerous capabilities" in order to…

Computation and Language · Computer Science 2023-09-06 Yuxia Wang , Haonan Li , Xudong Han , Preslav Nakov , Timothy Baldwin

Text-based safety classifiers are widely used for content moderation and increasingly to tune generative language model behavior - a topic of growing concern for the safety of digital assistants and chatbots. However, different policies…

Computation and Language · Computer Science 2023-10-24 Maximilian Mozes , Jessica Hoffmann , Katrin Tomanek , Muhamed Kouate , Nithum Thain , Ann Yuan , Tolga Bolukbasi , Lucas Dixon

Recent work introduced the model of learning from discriminative feature feedback, in which a human annotator not only provides labels of instances, but also identifies discriminative features that highlight important differences between…

Machine Learning · Computer Science 2021-05-25 Sanjoy Dasgupta , Sivan Sabato

The prevalence and impact of toxic discussions online have made content moderation crucial.Automated systems can play a vital role in identifying toxicity, and reducing the reliance on human moderation.Nevertheless, identifying toxic…

Artificial Intelligence · Computer Science 2023-11-02 Senjuti Dutta , Sid Mittal , Sherol Chen , Deepak Ramachandran , Ravi Rajakumar , Ian Kivlichan , Sunny Mak , Alena Butryna , Praveen Paritosh

Understanding human driving behavior is crucial to develop autonomous vehicles' algorithms. However, most low level automation, such as the one in advanced driving assistance systems (ADAS), is based on objective safety measures, which are…

Robotics · Computer Science 2022-11-03 Enrico Del Re , Cristina Olaverri-Monreal

Safety and responsibility evaluations of advanced AI models are a critical but developing field of research and practice. In the development of Google DeepMind's advanced AI models, we innovated on and applied a broad set of approaches to…

Ensuring safety in human-robot interaction (HRI) is essential to foster user trust and enable the broader adoption of robotic systems. Traditional safety models primarily rely on sensor-based measures, such as relative distance and…

Robotics · Computer Science 2025-07-10 Pranav Pandey , Ramviyas Parasuraman , Prashant Doshi

AI Safety Moderation (ASM) classifiers are designed to moderate content on social media platforms and to serve as guardrails that prevent Large Language Models (LLMs) from being fine-tuned on unsafe inputs. Owing to their potential for…

Computation and Language · Computer Science 2025-01-24 Akshit Achara , Anshuman Chhabra

As large language models (LLMs) become increasingly integrated into real-world applications, scalable and rigorous safety evaluation is essential. This paper introduces Aymara AI, a programmatic platform for generating and administering…

Artificial Intelligence · Computer Science 2026-05-01 Juan Manuel Contreras

We measure the association between generative AI (GAI) tool adoption and four metrics spanning security operations, information protection, and endpoint management: 1) number of security alerts per incident, 2) probability of security…

Cryptography and Security · Computer Science 2025-04-15 James Bono , Justin Grana , Kleanthis Karakolios , Pruthvi Hanumanthapura Ramakrishna , Ankit Srivastava

AI incident reporting requirements are emerging in regulation and policy, yet no operational criteria exist for determining when a detected AI incident warrants escalation beyond national handling to international coordination. This paper…

Computers and Society · Computer Science 2026-05-20 Francesca Gomez , Matthew Ball , Michael Harre , Lydia Preston , Josephine Schwab , Caio Machado

Modeling and calibrating the fidelity of synthetic data is paramount in shaping the future of safe and reliable self-driving technology by offering a cost-effective and scalable alternative to real-world data collection. We focus on its…

Software Engineering · Computer Science 2025-04-16 Chih-Hong Cheng , Paul Stöckel , Xingyu Zhao

Artificial intelligence (AI) systems are deployed as collaborators in human decision-making. Yet, evaluation practices focus primarily on model accuracy rather than whether human-AI teams are prepared to collaborate safely and effectively.…

Human-Computer Interaction · Computer Science 2026-03-20 Min Hun Lee

Natural Language Inference (NLI) is foundational for evaluating language understanding in AI. However, progress has plateaued, with models failing on ambiguous examples and exhibiting poor generalization. We argue that this stems from…

Computation and Language · Computer Science 2024-05-21 Claudiu Creanga , Liviu P. Dinu

In this work, we introduce a novel metric for auditing group fairness in ranked lists. Our approach offers two benefits compared to the state of the art. First, we offer a blueprint for modeling of user attention. Rather than assuming a…

Computers and Society · Computer Science 2019-05-14 Piotr Sapiezynski , Wesley Zeng , Ronald E. Robertson , Alan Mislove , Christo Wilson

Security evaluations inherently depend on stable identifiers. Any finding, audit, or regulatory decision must remain attached to the specific artifact it pertains to. Continuously updated artificial intelligence systems violate this core…

Cryptography and Security · Computer Science 2026-05-26 Dan Ristea , Vasilios Mavroudis

In commonsense generation, given a set of input concepts, a model must generate a response that is not only commonsense bearing, but also capturing multiple diverse viewpoints. Numerous evaluation metrics based on form- and content-level…

Computation and Language · Computer Science 2025-06-03 Tianhui Zhang , Bei Peng , Danushka Bollegala