English

Evaluation of Large Language Models for Numeric Anomaly Detection in Power Systems

Systems and Control 2025-12-09 v2 Systems and Control

Abstract

Large language models (LLMs) have gained increasing attention in power grids for their general-purpose capabilities. Meanwhile, anomaly detection (AD) remains critical for grid resilience, requiring accurate and interpretable decisions based on multivariate telemetry. Yet the performance of LLMs on large-scale numeric data for AD remains largely unexplored. This paper presents a comprehensive evaluation of LLMs for numeric AD in power systems. We use GPT-OSS-20B as a representative model and evaluate it on the IEEE 14-bus system. A standardized prompt framework is applied across zero-shot, few-shot, in-context learning, low rank adaptation (LoRA), fine-tuning, and a hybrid LLM-traditional approach. We adopt a rule-aware design based on the three-sigma criterion, and report detection performance and rationale quality. This study lays the groundwork for further investigation into the limitations and capabilities of LLM-based AD and its integration with classical detectors in cyber-physical power grid applications.

Keywords

Cite

@article{arxiv.2511.21371,
  title  = {Evaluation of Large Language Models for Numeric Anomaly Detection in Power Systems},
  author = {Yichen Liu and Hongyu Wu and Bo Liu},
  journal= {arXiv preprint arXiv:2511.21371},
  year   = {2025}
}
R2 v1 2026-07-01T07:56:10.316Z