English

ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders

Computation and Language 2026-02-20 v1

Abstract

The promise of LLM-based user simulators to improve conversational AI is hindered by a critical "realism gap," leading to systems that are optimized for simulated interactions, but may fail to perform well in the real world. We introduce ConvApparel, a new dataset of human-AI conversations designed to address this gap. Its unique dual-agent data collection protocol -- using both "good" and "bad" recommenders -- enables counterfactual validation by capturing a wide spectrum of user experiences, enriched with first-person annotations of user satisfaction. We propose a comprehensive validation framework that combines statistical alignment, a human-likeness score, and counterfactual validation to test for generalization. Our experiments reveal a significant realism gap across all simulators. However, the framework also shows that data-driven simulators outperform a prompted baseline, particularly in counterfactual validation where they adapt more realistically to unseen behaviors, suggesting they embody more robust, if imperfect, user models.

Keywords

Cite

@article{arxiv.2602.16938,
  title  = {ConvApparel: A Benchmark Dataset and Validation Framework for User Simulators in Conversational Recommenders},
  author = {Ofer Meshi and Krisztian Balog and Sally Goldman and Avi Caciularu and Guy Tennenholtz and Jihwan Jeong and Amir Globerson and Craig Boutilier},
  journal= {arXiv preprint arXiv:2602.16938},
  year   = {2026}
}

Comments

EACL 2026

R2 v1 2026-07-01T10:42:14.090Z