English

Can AI Recognize Its Own Reflection? Self-Detection Performance of LLMs in Computing Education

Computers and Society 2025-12-30 v1

Abstract

The rapid advancement of Large Language Models (LLMs) presents a significant challenge to academic integrity within computing education. As educators seek reliable detection methods, this paper evaluates the capacity of three prominent LLMs (GPT-4, Claude, and Gemini) to identify AI-generated text in computing-specific contexts. We test their performance under both standard and 'deceptive' prompt conditions, where the models were instructed to evade detection. Our findings reveal a significant instability: while default AI-generated text was easily identified, all models struggled to correctly classify human-written work (with error rates up to 32%). Furthermore, the models were highly susceptible to deceptive prompts, with Gemini's output completely fooling GPT-4. Given that simple prompt alterations significantly degrade detection efficacy, our results demonstrate that these LLMs are currently too unreliable for making high-stakes academic misconduct judgments.

Keywords

Cite

@article{arxiv.2512.23587,
  title  = {Can AI Recognize Its Own Reflection? Self-Detection Performance of LLMs in Computing Education},
  author = {Christopher Burger and Karmece Talley and Christina Trotter},
  journal= {arXiv preprint arXiv:2512.23587},
  year   = {2025}
}

Comments

10 pages, 5 tables. Accepted for publication at the 59th Hawaii International Conference on System Sciences

R2 v1 2026-07-01T08:44:34.633Z