Artificial Intelligence · Computer Science
AgentMisalignment: Measuring the Propensity for Misaligned Behaviour in LLM-Based Agents
Akshat Naik, Patrick Quinn, Guillermo Bosch, Emma Gouné +3
2025-10-02
Computation and Language · Computer Science
LLMs Deceive Unintentionally: Emergent Misalignment in Dishonesty from Misaligned Samples to Biased Human-AI Interactions
Xuhao Hu, Peng Wang, Xiaoya Lu, Dongrui Liu +2
2026-01-21
Machine Learning · Computer Science
Coercing LLMs to do and reveal (almost) anything
Jonas Geiping, Alex Stein, Manli Shu, Khalid Saifullah +2
2024-02-22
Machine Learning · Computer Science
Thought Crime: Backdoors and Emergent Misalignment in Reasoning Models
James Chua, Jan Betley, Mia Taylor, Owain Evans
2025-07-11
Machine Learning · Computer Science
Can LLMs Lie? Investigation beyond Hallucination
Haoran Huan, Mihir Prabhudesai, Mengning Wu, Shantanu Jaiswal +1
2025-09-04
Computation and Language · Computer Science
How Alignment and Jailbreak Work: Explain LLM Safety through Intermediate Hidden States
Zhenhong Zhou, Haiyang Yu, Xinghua Zhang, Rongwu Xu +2
2024-06-14
Cryptography and Security · Computer Science
Agentic Misalignment: How LLMs Could Be Insider Threats
Aengus Lynch, Benjamin Wright, Caleb Larson, Stuart J. Ritchie +4
2025-10-17
Artificial Intelligence · Computer Science
When Do LLM Preferences Predict Downstream Behavior?
Katarina Slama, Alexandra Souly, Dishank Bansal, Henry Davidson +2
2026-02-24
Computation and Language · Computer Science
Emergent Misalignment: Narrow finetuning can produce broadly misaligned LLMs
Jan Betley, Daniel Tan, Niels Warncke, Anna Sztyber-Betley +4
2026-01-27
Cryptography and Security · Computer Science
Confusion is the Final Barrier: Rethinking Jailbreak Evaluation and Investigating the Real Misuse Threat of LLMs
Yu Yan, Sheng Sun, Zhe Wang, Yijun Lin +5
2025-09-16
Artificial Intelligence · Computer Science
LLM Censorship: A Machine Learning Challenge or a Computer Security Problem?
David Glukhov, Ilia Shumailov, Yarin Gal, Nicolas Papernot +1
2023-07-25
Computers and Society · Computer Science
Chat Bankman-Fried: an Exploration of LLM Alignment in Finance
Claudia Biancotti, Carolina Camassa, Andrea Coletta, Oliver Giudice +1
2025-02-26
Computation and Language · Computer Science
Uncovering the Persuasive Fingerprint of LLMs in Jailbreaking Attacks
Havva Alizadeh Noughabi, Julien Serbanescu, Fattane Zarrinkalam, Ali Dehghantanha
2025-10-28
Computation and Language · Computer Science
Competition Report: Finding Universal Jailbreak Backdoors in Aligned LLMs
Javier Rando, Francesco Croce, Kryštof Mitka, Stepan Shabalin +3
2024-06-07
Computation and Language · Computer Science
A Course Correction in Steerability Evaluation: Revealing Miscalibration and Side Effects in LLMs
Trenton Chang, Tobias Schnabel, Adith Swaminathan, Jenna Wiens
2026-01-21
Machine Learning · Computer Science
Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
Alexander Panfilov, Evgenii Kortukov, Kristina Nikolić, Matthias Bethge +5
2025-09-24