Graph neural network-based network intrusion detection systems have recently demonstrated state-of-the-art performance on benchmark datasets. Nevertheless, these methods suffer from a reliance on target encoding for data pre-processing, limiting widespread adoption due to the associated need for annotated labels--a cost-prohibitive requirement. In this work, we propose a solution involving in-context pre-training and the utilization of dense representations for categorical features to jointly overcome the label-dependency limitation. Our approach exhibits remarkable data efficiency, achieving over 98% of the performance of the supervised state-of-the-art with less than 4% labeled data on the NF-UQ-NIDS-V2 dataset.
@article{arxiv.2402.18986,
title = {Always be Pre-Training: Representation Learning for Network Intrusion Detection with GNNs},
author = {Zhengyao Gu and Diego Troy Lopez and Lilas Alrahis and Ozgur Sinanoglu},
journal= {arXiv preprint arXiv:2402.18986},
year = {2024}
}
Comments
Will appear in the 2024 International Symposium on Quality Electronic Design (ISQED'24)