English

Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach

Computers and Society 2026-01-21 v2

Abstract

Current safety alignment for Large Language Models (LLMs) implicitly optimizes for a "modal adult user," leaving models vulnerable to distributional shifts in user cognition. We present ChildSafe, a benchmark that quantifies alignment robustness under cognitive shifts corresponding to four developmental stages. Unlike static persona-based evaluations, we introduce a parametric cognitive simulation approach, formalizing developmental stages as hyperparameter constraints (e.g., volatility, context horizon) to generate out-of-distribution interaction traces. We validate these agents against ground-truth human linguistic data (CHILDES) and deploy them across 1,200 multi-turn interactions. Our results reveal a systematic alignment generalization gap: state-of-the-art models exhibit up to 11.5% performance degradation when interacting with early-childhood agents compared to standard baselines. We provide the research community with the validated agent artifacts and evaluation protocols to facilitate robust alignment testing against non-adversarial, cognitively diverse populations.

Keywords

Cite

@article{arxiv.2510.05484,
  title  = {Evaluating LLM Safety Across Child Development Stages: A Simulated Agent Approach},
  author = {Abhejay Murali and Saleh Afroogh and Kevin Chen and David Atkinson and Amit Dhurandhar and Junfeng Jiao},
  journal= {arXiv preprint arXiv:2510.05484},
  year   = {2026}
}
R2 v1 2026-07-01T06:20:24.493Z