English

Security of Language Models for Code: A Systematic Literature Review

Software Engineering 2025-05-20 v2 Cryptography and Security

Abstract

Language models for code (CodeLMs) have emerged as powerful tools for code-related tasks, outperforming traditional methods and standard machine learning approaches. However, these models are susceptible to security vulnerabilities, drawing increasing research attention from domains such as software engineering, artificial intelligence, and cybersecurity. Despite the growing body of research focused on the security of CodeLMs, a comprehensive survey in this area remains absent. To address this gap, we systematically review 67 relevant papers, organizing them based on attack and defense strategies. Furthermore, we provide an overview of commonly used language models, datasets, and evaluation metrics, and highlight open-source tools and promising directions for future research in securing CodeLMs.

Keywords

Cite

@article{arxiv.2410.15631,
  title  = {Security of Language Models for Code: A Systematic Literature Review},
  author = {Yuchen Chen and Weisong Sun and Chunrong Fang and Zhenpeng Chen and Yifei Ge and Tingxu Han and Quanjun Zhang and Yang Liu and Zhenyu Chen and Baowen Xu},
  journal= {arXiv preprint arXiv:2410.15631},
  year   = {2025}
}

Comments

Accepted to ACM Transactions on Software Engineering and Methodology (TOSEM)

R2 v1 2026-06-28T19:29:06.446Z