English
Related papers

Related papers: Attention Shift: Steering AI Away from Unsafe Cont…

200 papers

The evolution of large language models into autonomous agents introduces adversarial failures that exploit legitimate tool privileges, transforming safety evaluation in tool-augmented environments from a subjective NLP task into an…

Machine Learning · Computer Science 2026-02-03 Samuel Nellessen , Tal Kachman

Toxicity detection algorithms, originally designed with reactive content moderation in mind, are increasingly being deployed into proactive end-user interventions to moderate content. Through a socio-technical lens and focusing on contexts…

Human-Computer Interaction · Computer Science 2025-02-25 Mark Warner , Angelika Strohmayer , Matthew Higgs , Lynne Coventry

Generative AI, with its tendency to "hallucinate" incorrect results, may pose a risk to knowledge work by introducing errors. On the other hand, it may also provide unprecedented opportunities for users, particularly non-experts, to learn…

Human-Computer Interaction · Computer Science 2024-12-20 Advait Sarkar , Xiaotong , Xu , Neil Toronto , Ian Drosos , Christian Poelitz

Dynamic learning systems subject to selective labeling exhibit censoring, i.e. persistent negative predictions assigned to one or more subgroups of points. In applications like consumer finance, this results in groups of applicants that are…

Machine Learning · Computer Science 2023-06-30 Jennifer Chien , Margaret Roberts , Berk Ustun

Artificial intelligence generated content (AIGC), a rapidly advancing technology, is transforming content creation across domains, such as text, images, audio, and video. Its growing potential has attracted more and more researchers and…

Artificial Intelligence · Computer Science 2026-03-03 Chengzhang Zhu , Luobin Cui , Ying Tang , Jiacun Wang

Recent computational approaches for combating online hate speech involve the automatic generation of counter narratives by adapting Pretrained Transformer-based Language Models (PLMs) with human-curated data. This process, however, can…

Computation and Language · Computer Science 2023-09-06 Helena Bonaldi , Giuseppe Attanasio , Debora Nozza , Marco Guerini

Large language models increasingly rely on explicit chain-of-thought reasoning to solve complex tasks, yet the safety of the reasoning process itself remains largely unaddressed. Existing work focuses predominantly on content safety (i.e.,…

Artificial Intelligence · Computer Science 2026-05-07 Xunguang Wang , Yuguang Zhou , Qingyue Wang , Zongjie Li , Ruixuan Huang , Zhenlan Ji , Pingchuan Ma , Shuai Wang

Generative AI poses both opportunities and risks for solving inverse design problems in the sciences. Generative tools provide the ability to expand and refine a search space autonomously, but do so at the cost of exploring low-quality…

Machine Learning · Computer Science 2025-10-01 Marcus Schwarting , Logan Ward , Nathaniel Hudson , Xiaoli Yan , Ben Blaiszik , Santanu Chaudhuri , Eliu Huerta , Ian Foster

Social media platforms struggle to protect users from harmful content through content moderation. These platforms have recently leveraged machine learning models to cope with the vast amount of user-generated content daily. Since moderation…

Machine Learning · Computer Science 2023-01-27 Donghyun Son , Byounggyu Lew , Kwanghee Choi , Yongsu Baek , Seungwoo Choi , Beomjun Shin , Sungjoo Ha , Buru Chang

AI-generated content (AIGC) detectors are increasingly deployed in high-stakes settings such as academic integrity screening, yet their reliability rests on a fundamental paradox: as language models are trained on human-written corpora, the…

Machine Learning · Computer Science 2026-05-05 Guantian Zheng

Generative AI models offer powerful capabilities but often lack transparency, making it difficult to interpret their output. This is critical in cases involving artistic or copyrighted content. This work introduces a search-inspired…

Artificial Intelligence · Computer Science 2025-04-03 Theodoros Aivalis , Iraklis A. Klampanos , Antonis Troumpoukis , Joemon M. Jose

We propose a new approach to promote safety in classification tasks with established concepts. Our approach -- called a conceptual safeguard -- acts as a verification layer for models that predict a target outcome by first predicting the…

Machine Learning · Computer Science 2024-11-08 Hailey Joren , Charles Marx , Berk Ustun

Adding additional control to pretrained diffusion models has become an increasingly popular research area, with extensive applications in computer vision, reinforcement learning, and AI for science. Recently, several studies have proposed…

Machine Learning · Computer Science 2024-05-30 Yifei Shen , Xinyang Jiang , Yezhen Wang , Yifan Yang , Dongqi Han , Dongsheng Li

We introduce SAFEMax, a novel method for Machine Unlearning in diffusion models. Grounded in information-theoretic principles, SAFEMax maximizes the entropy in generated images, causing the model to generate Gaussian noise when conditioned…

Machine Learning · Computer Science 2025-08-29 Christoforos N. Spartalis , Theodoros Semertzidis , Petros Daras , Efstratios Gavves

Trajectory prediction models in autonomous driving are vulnerable to perturbations from non-causal agents whose actions should not affect the ego-agent's behavior. Such perturbations can lead to incorrect predictions of other agents'…

Robotics · Computer Science 2026-05-19 Ehsan Ahmadi , Ray Mercurius , Soheil Alizadeh , Kasra Rezaee , Amir Rasouli

As generative AI systems, including large language models (LLMs) and diffusion models, advance rapidly, their growing adoption has led to new and complex security risks often overlooked in traditional AI risk assessment frameworks. This…

Cryptography and Security · Computer Science 2024-10-21 Aviral Srivastava , Sourav Panda

We propose a novel interactive learning framework which we refer to as Interactive Attention Learning (IAL), in which the human supervisors interactively manipulate the allocated attentions, to correct the model's behavior by updating the…

Machine Learning · Computer Science 2020-06-11 Jay Heo , Junhyeon Park , Hyewon Jeong , Kwang Joon Kim , Juho Lee , Eunho Yang , Sung Ju Hwang

Recent advances in text-to-image generative models have raised concerns about their potential to produce harmful content when provided with malicious input text prompts. To address this issue, two main approaches have emerged: (1)…

Machine Learning · Computer Science 2025-11-13 Jiwoo Shin , Byeonghu Na , Mina Kang , Wonhyeok Choi , Il-Chul Moon

High-risk industries like nuclear and aviation use real-time monitoring to detect dangerous system conditions. Similarly, Large Language Models (LLMs) need monitoring safeguards. We propose a real-time framework to predict harmful AI…

Artificial Intelligence · Computer Science 2025-05-21 Maheep Chaudhary , Fazl Barez

When deployed in the real world, machine learning models inevitably encounter changes in the data distribution, and certain -- but not all -- distribution shifts could result in significant performance degradation. In practice, it may make…

Machine Learning · Statistics 2022-05-06 Aleksandr Podkopaev , Aaditya Ramdas