English

Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data

Computation and Language 2026-06-27 v1 Computers and Society Machine Learning

Abstract

Large language models (LLMs) are increasingly used to simulate social survey responses, yet their outputs exhibit systematic biases: marginal distributions are skewed, response variance is poorly calibrated, and predictor-outcome relationships are attenuated. We ask a simple question: given a small pilot sample of human responses, can an LLM recover the statistical characteristics of a broader population? We decompose recovery along three axes: structural fidelity, marginal fidelity, and individual fidelity. Using a COVID-19 misinformation survey as a case study, we benchmark three families of approaches: prompting, rectification, and fine-tuning. The findings suggest that fine-tuning on small pilot samples offers a balanced approach for achieving multiple forms of fidelity, but the levels of such fidelity can vary across subsamples, potentially threatening pluralistic alignment.

Cite

@article{arxiv.2606.28963,
  title  = {Beyond the Mean: Three-Axis Fidelity for Aligning LLM-Based Survey Simulators from Small Pilot Data},
  author = {Eun Cheol Choi and Youngrae Kim and Prabhu Pugalenthi and Hong-En Chen and Bo-Ruei Huang},
  journal= {arXiv preprint arXiv:2606.28963},
  year   = {2026}
}

Comments

11 pages, 8 tables, 3 figures; Pluralistic Alignment @ ICML 2026 Workshop