English

Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users

Artificial Intelligence 2025-10-22 v2 Computation and Language

Abstract

We study a web-deployed, tool-augmented LLM health coach with real users. In a pilot with seven users (280 rated turns), offline policy evaluation (OPE) over factorized decision heads (Tool/Style) shows that a uniform heavy-tool policy raises average value on logs but harms specific subgroups, most notably low-health-literacy/high-self-efficacy users. A lightweight simulator with hidden archetypes further shows that adding a small early information-gain bonus reliably shortens trait identification and improves goal success and pass@3. Together, these early findings indicate an evaluation-first path to personalization: freeze the generator, learn subgroup-aware decision heads on typed rewards (objective tool outcomes and satisfaction), and always report per-archetype metrics to surface subgroup harms that averages obscure.

Keywords

Cite

@article{arxiv.2510.17173,
  title  = {Offline Policy Evaluation of Multi-Turn LLM Health Coaching with Real Users},
  author = {Melik Ozolcer and Sang Won Bae},
  journal= {arXiv preprint arXiv:2510.17173},
  year   = {2025}
}

Comments

Accepted to the NeurIPS 2025 Workshop on Multi-Turn Interactions in Large Language Models