English

StagePilot: A Deep Reinforcement Learning Agent for Stage-Controlled Cybergrooming Simulation

Machine Learning 2026-02-06 v1 Computation and Language

Abstract

Cybergrooming is an evolving threat to youth, necessitating proactive educational interventions. We propose StagePilot, an offline RL-based dialogue agent that simulates the stage-wise progression of grooming behaviors for prevention training. StagePilot selects conversational stages using a composite reward that balances user sentiment and goal proximity, with transitions constrained to adjacent stages for realism and interpretability. We evaluate StagePilot through LLM-based simulations, measuring stage completion, dialogue efficiency, and emotional engagement. Results show that StagePilot generates realistic and coherent conversations aligned with grooming dynamics. Among tested methods, the IQL+AWAC agent achieves the best balance between strategic planning and emotional coherence, reaching the final stage up to 43% more frequently than baselines while maintaining over 70% sentiment alignment.

Keywords

Cite

@article{arxiv.2602.05060,
  title  = {StagePilot: A Deep Reinforcement Learning Agent for Stage-Controlled Cybergrooming Simulation},
  author = {Heajun An and Qi Zhang and Minqian Liu and Xinyi Zhang and Sang Won Lee and Lifu Huang and Pamela J. Wisniewski and Jin-Hee Cho},
  journal= {arXiv preprint arXiv:2602.05060},
  year   = {2026}
}
R2 v1 2026-07-01T09:36:50.076Z