Computer Vision and Pattern Recognition · Computer Science
Safety Without Semantic Disruptions: Editing-free Safe Image Generation via Context-preserving Dual Latent Reconstruction
Jordan Vice, Naveed Akhtar, Mubarak Shah, Richard Hartley +1
2025-03-06
Machine Learning · Computer Science
An Adversarial Perspective on Machine Unlearning for AI Safety
Jakub Łucki, Boyi Wei, Yangsibo Huang, Peter Henderson +2
2025-06-03
Machine Learning · Computer Science
Distributional Surgery for Language Model Activations
Bao Nguyen, Binh Nguyen, Duy Nguyen, Viet Anh Nguyen
2025-11-11
Artificial Intelligence · Computer Science
From Refusal to Recovery: A Control-Theoretic Approach to Generative AI Guardrails
Ravi Pandya, Madison Bland, Duy P. Nguyen, Changliu Liu +2
2026-05-20
Machine Learning · Computer Science
Machine Unlearning Doesn't Do What You Think: Lessons for Generative AI Policy and Research
A. Feder Cooper, Christopher A. Choquette-Choo, Miranda Bogen, Kevin Klyman +33
2026-02-26
Computer Vision and Pattern Recognition · Computer Science
To Generate or Not? Safety-Driven Unlearned Diffusion Models Are Still Easy To Generate Unsafe Images ... For Now
Yimeng Zhang, Jinghan Jia, Xin Chen, Aochuan Chen +4
2024-07-09
Computer Vision and Pattern Recognition · Computer Science
SAFREE: Training-Free and Adaptive Guard for Safe Text-to-Image And Video Generation
Jaehong Yoon, Shoubin Yu, Vaidehi Patil, Huaxiu Yao +1
2025-03-17
Machine Learning · Computer Science
Safety Pretraining: Toward the Next Generation of Safe AI
Pratyush Maini, Sachin Goyal, Dylan Sam, Alex Robey +6
2025-09-16
Machine Learning · Computer Science
Backtracking Improves Generation Safety
Yiming Zhang, Jianfeng Chi, Hailey Nguyen, Kartikeya Upasani +3
2024-09-24
Cryptography and Security · Computer Science
From AI-Generated Content to Agentic Action: Security and Safety Threats in Generative AI
Zelin Zhang, Qi Li, Jie Cao, Lingshuang Liu +1
2026-05-19
Computation and Language · Computer Science
Wisdom is Knowing What not to Say: Hallucination-Free LLMs Unlearning via Attention Shifting
Chenchen Tan, Youyang Qu, Xinghao Li, Hui Zhang +3
2026-04-20
Computer Vision and Pattern Recognition · Computer Science
Video Unlearning via Low-Rank Refusal Vector
Simone Facchiano, Stefano Saravalle, Matteo Migliarini, Edoardo De Matteis +6
2026-02-02
Machine Learning · Computer Science
Safety and Fairness for Content Moderation in Generative Models
Susan Hao, Piyush Kumar, Sarah Laszlo, Shivani Poddar +2
2023-06-13
Cryptography and Security · Computer Science
Attention Eclipse: Manipulating Attention to Bypass LLM Safety-Alignment
Pedram Zaree, Md Abdullah Al Mamun, Quazi Mishkatul Alam, Yue Dong +2
2025-02-24
Computer Vision and Pattern Recognition · Computer Science
SafeDiffusion-R1: Online Reward Steering for Safe Diffusion Post-Training
Komal Kumar, Ankan Deria, Abhishek Basu, Fahad Shamshad +2
2026-05-19
Machine Learning · Computer Science
Mitigating Privacy Risk via Forget Set-Free Unlearning
Aviraj Newatia, Michael Cooper, Viet Nguyen, Rahul G. Krishnan
2026-04-14