Machine Learning · Computer Science
Subversion Strategy Eval: Can language models statelessly strategize to subvert control protocols?
Alex Mallen, Charlie Griffin, Misha Wagner, Alessandro Abate +1
2025-04-07
Software Engineering · Computer Science
Hierarchical Evaluation of Software Design Capabilities of Large Language Models of Code
Mootez Saad, Boqi Chen, José Antonio Hernández López, Dániel Varró +1
2025-12-30
Artificial Intelligence · Computer Science
Concurrent Linguistic Error Detection (CLED): a New Methodology for Error Detection in Large Language Models
Jinhua Zhu, Javier Conde, Zhen Gao, Pedro Reviriego +2
2025-09-17
Computation and Language · Computer Science
Where Do Backdoors Live? A Component-Level Analysis of Backdoor Propagation in Speech Language Models
Alexandrine Fortier, Thomas Thebaud, Jesús Villalba, Najim Dehak +2
2026-04-08
Machine Learning · Computer Science
AI Control: Improving Safety Despite Intentional Subversion
Ryan Greenblatt, Buck Shlegeris, Kshitij Sachan, Fabien Roger
2024-07-24
Computation and Language · Computer Science
Among Us: Measuring and Mitigating Malicious Contributions in Model Collaboration Systems
Ziyuan Yang, Wenxuan Ding, Shangbin Feng, Yulia Tsvetkov
2026-02-06
Cryptography and Security · Computer Science
A Survey of Recent Backdoor Attacks and Defenses in Large Language Models
Shuai Zhao, Meihuizi Jia, Zhongliang Guo, Leilei Gan +6
2025-01-07
Cryptography and Security · Computer Science
Misusing Tools in Large Language Models With Visual Adversarial Examples
Xiaohan Fu, Zihan Wang, Shuheng Li, Rajesh K. Gupta +3
2023-10-06
Computation and Language · Computer Science
False Sense of Security: Why Probing-based Malicious Input Detection Fails to Generalize
Cheng Wang, Zeming Wei, Qin Liu, Muhao Chen
2025-12-16
Software Engineering · Computer Science
Automatic Detection of LLM-Generated Code: A Comparative Case Study of Contemporary Models Across Function and Class Granularities
Musfiqur Rahman, SayedHassan Khatoonabadi, Ahmad Abdellatif, Emad Shihab
2025-12-23
Computation and Language · Computer Science
Evaluating Machine Common Sense via Cloze Testing
Ehsan Qasemi, Lee Kezar, Jay Pujara, Pedro Szekely
2022-01-21
Cryptography and Security · Computer Science
ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety
Kun Wang, Cheng Qian, Miao Yu, Lilan Peng +5
2026-04-22
Software Engineering · Computer Science
The Struggles of LLMs in Cross-lingual Code Clone Detection
Micheline Bénédicte Moumoula, Abdoul Kader Kabore, Jacques Klein, Tegawendé Bissyande
2025-05-07
Computation and Language · Computer Science
Chain-of-Code Collapse: Reasoning Failures in LLMs via Adversarial Prompting in Code Generation
Jaechul Roh, Varun Gandhi, Shivani Anilkumar, Arin Garg
2025-06-13
Artificial Intelligence · Computer Science
Exploring the Adversarial Capabilities of Large Language Models
Lukas Struppek, Minh Hieu Le, Dominik Hintersdorf, Kristian Kersting
2024-07-09
Software Engineering · Computer Science
Programming Language Confusion: When Code LLMs Can't Keep their Languages Straight
Micheline Bénédicte Moumoula, Serge Lionel Nikiema, Abdoul Kader Kabore, Jacques Klein +1
2026-02-03
Cryptography and Security · Computer Science
Prompt Injection Attacks on Large Language Models in Oncology
Jan Clusmann, Dyke Ferber, Isabella C. Wiest, Carolin V. Schneider +4
2025-03-20