English
Related papers

Related papers: An Empirical Evaluation of LLMs for Solving Offens…

200 papers

The assessment of cybersecurity Capture-The-Flag (CTF) exercises involves participants finding text strings or ``flags'' by exploiting system vulnerabilities. Large Language Models (LLMs) are natural-language models trained on vast amounts…

Artificial Intelligence · Computer Science 2023-08-22 Wesley Tann , Yuancheng Liu , Jun Heng Sim , Choon Meng Seah , Ee-Chien Chang

Large Language Models (LLMs) are being deployed across various domains today. However, their capacity to solve Capture the Flag (CTF) challenges in cybersecurity has not been thoroughly evaluated. To address this, we develop a novel method…

Capture-the-Flag (CTF) competitions are crucial for cybersecurity education and training. As large language models (LLMs) evolve, there is increasing interest in their ability to automate CTF challenge solving. For example, DARPA has…

Artificial Intelligence · Computer Science 2025-06-24 Zimo Ji , Daoyuan Wu , Wenyuan Jiang , Pingchuan Ma , Zongjie Li , Shuai Wang

Large Language Model (LLM) agents are increasingly proposed to automate offensive security tasks, with recent studies reporting near human-level success rates in Capture-the-Flag (CTF) challenges. We here revisit these results, providing a…

Cryptography and Security · Computer Science 2026-05-22 Youness Bouchari , Matteo Boffa , Marco Mellia , Idilio Drago , Thanh Minh Bui , Dario Rossi

Team collaboration among individuals with diverse sets of expertise and skills is essential for solving complex problems. As part of an interdisciplinary effort, we studied the effects of Capture the Flag (CTF) game, a popular and engaging…

Computers and Society · Computer Science 2022-06-22 Sang-Yoon Chang , Kay Yoon , Simeon Wuthier , Kelei Zhang

Cybersecurity training has become a crucial part of computer science education and industrial onboarding. Capture the Flag (CTF) competitions have emerged as a valuable, gamified approach for developing and refining the skills of…

Cryptography and Security · Computer Science 2025-06-02 Yuwei Yang , Skyler Grandel , Daniel Balasubramanian , Yu Huang , Kevin Leach

Recent advances in LLM agentic systems have improved the automation of offensive security tasks, particularly for Capture the Flag (CTF) challenges. We systematically investigate the key factors that drive agent success and provide a…

Agentic large language models (LLMs) are increasingly evaluated on cybersecurity tasks using capture-the-flag (CTF) benchmarks, yet existing pointwise benchmarks offer limited insight into agent robustness and generalisation across…

Software Engineering · Computer Science 2026-04-20 Shahin Honarvar , Amber Gorzynski , James Lee-Jones , Harry Coppock , Marek Rei , Joseph Ryan , Alastair F. Donaldson

Capture-the-Flag (CTF) competitions are increasingly becoming a testbed for evaluating AI capabilities at solving security tasks, due to the controlled environments and objective success criteria. Existing evaluations have focused on how…

Cryptography and Security · Computer Science 2026-02-25 Tingxuan Tang , Nicolas Janis , Kalyn Asher Montague , Kevin Eykholt , Dhilung Kirat , Youngja Park , Jiyong Jang , Adwait Nadkarni , Yue Xiao

Capture the Flag challenges are a popular form of cybersecurity education, where students solve hands-on tasks in an informal, game-like setting. The tasks feature diverse assignments, such as exploiting websites, cracking passwords, and…

Cryptography and Security · Computer Science 2021-01-06 Valdemar Švábenský , Pavel Čeleda , Jan Vykopal , Silvia Brišáková

Capture-the-Flag (CTF) competitions play a central role in modern cybersecurity as a platform for training practitioners and evaluating offensive and defensive techniques derived from real-world vulnerabilities. Despite recent advances in…

Cryptography and Security · Computer Science 2026-01-15 Xiaonan Liu , Zhihao Li , Xiao Lan , Hao Ren , Haizhou Wang , Xingshu Chen

Large Language Models (LLMs) have been used in cybersecurity such as autonomous security analysis or penetration testing. Capture the Flag (CTF) challenges serve as benchmarks to assess automated task-planning abilities of LLM agents for…

Capture the Flag (CTF) competitions represent a powerful experiential learning approach within cybersecurity education, blending diverse concepts into interactive challenges. However, the short duration (typically 24-48 hours) and ephemeral…

Cryptography and Security · Computer Science 2025-12-02 Pratham Gupta , Aditya Gabani , Connor Nelson , Yan Shoshitaishvili

Recent advances in Large Language Models (LLMs) have enabled agentic systems for complex, multi-step tasks; cybersecurity is emerging as a prominent application. To evaluate such agents, researchers widely adopt Capture The Flag (CTF)…

Machine Learning · Computer Science 2026-05-13 Dongjun Lee , Ga-eun Bae , Insu Yun

Large Language Model (LLM) agents are increasingly proposed for autonomous cybersecurity tasks, but their capabilities in realistic offensive settings remain poorly understood. We present DeepRed, an open-source benchmark for evaluating…

Artificial Intelligence · Computer Science 2026-05-07 Ali Al-Kaswan , Maksim Plotnikov , Maxim Hájek , Roland Vízner , Arie van Deursen , Maliheh Izadi

In cybersecurity, Intrusion Detection Systems (IDS) serve as a vital defensive layer against adversarial threats. Accurate benchmarking is critical to evaluate and improve IDS effectiveness, yet traditional methodologies face limitations…

Cryptography and Security · Computer Science 2025-01-22 Manuel Kern , Florian Skopik , Max Landauer , Edgar Weippl

As large language models (LLMs) continue to evolve, their potential use in automating cyberattacks becomes increasingly likely. With capabilities such as reconnaissance, exploitation, and command execution, LLMs could soon become integral…

Cryptography and Security · Computer Science 2024-10-22 Daniel Ayzenshteyn , Roy Weiss , Yisroel Mirsky

We present 'Random-Crypto', a procedurally generated cryptographic Capture The Flag (CTF) dataset designed to unlock the potential of Reinforcement Learning (RL) for LLM-based agents in security-sensitive domains. Cryptographic reasoning…

Cryptography and Security · Computer Science 2025-08-19 Lajos Muzsai , David Imolai , András Lukács

Large language model systems face important security risks from maliciously crafted messages that aim to overwrite the system's original instructions or leak private data. To study this problem, we organized a capture-the-flag competition…

This study evaluates the ability of GPT-4o to autonomously solve beginner-level offensive security tasks by connecting the model to OverTheWire's Bandit capture-the-flag game. Of the 25 levels that were technically compatible with a…

Cryptography and Security · Computer Science 2026-01-27 Isabelle Bakker , John Hastings
‹ Prev 1 2 3 10 Next ›