English

SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use

Cryptography and Security 2026-05-26 v3 Artificial Intelligence

Abstract

Large Language Models (LLMs) are transforming enterprise workflows but introduce security and ethics challenges when employees inadvertently share confidential data or generate policy-violating content. This paper proposes SafeGPT, a two-sided guardrail system preventing sensitive data leakage and unethical outputs. SafeGPT integrates input-side detection/redaction, output-side moderation/reframing, and human-in-the-loop feedback. Experiments demonstrate SafeGPT effectively reduces data leakage risk and biased outputs while maintaining satisfaction.

Keywords

Cite

@article{arxiv.2601.06366,
  title  = {SafeGPT: Preventing Data Leakage and Unethical Outputs in Enterprise LLM Use},
  author = {Pratyush Desai and Luoxi Tang and Yuqiao Meng and Zhaohan Xi},
  journal= {arXiv preprint arXiv:2601.06366},
  year   = {2026}
}