English

From Tool to Teammate: LLM Coding Agents as Collaborative Partners for Behavioral Labeling in Educational Dialogue Analysis

Human-Computer Interaction 2026-03-31 v1

Abstract

Behavioral analysis of tutoring dialogues is essential for understanding student learning, yet manual coding remains a bottleneck. We present a methodology where LLM coding agents autonomously improve the prompts used by LLM classifiers to label educational dialogues. In each iteration, a coding agent runs the classifier against human-labeled validation data, analyzes disagreements, and proposes theory-grounded prompt modifications for researcher review. Applying this approach to 659 AI tutoring sessions across four experiments with three agents and three classifiers, 4-fold cross-validation on held-out data confirmed genuine improvement: the best agent achieved test κ=0.78\kappa=0.78 (SD=0.08=0.08), matching human inter-rater reliability (κ=0.78\kappa=0.78), at a cost of approximately $5--8 per agent. While development-set performance reached κ=0.91\kappa=0.91--0.930.93, the cross-validated results represent our primary generalization claim. The iterative process also surfaced an undocumented labeling pattern: human coders consistently treated expressions of confusion as engagement rather than disengagement. Continued iteration beyond the optimum led to regression, underscoring the need for held-out validation. We release all prompts, iteration logs, and data.

Keywords

Cite

@article{arxiv.2603.27440,
  title  = {From Tool to Teammate: LLM Coding Agents as Collaborative Partners for Behavioral Labeling in Educational Dialogue Analysis},
  author = {Eason Chen and Isabel Wang and Nina Yuan and Sophia Judicke and Kayla Beigh and Xinyi Tang},
  journal= {arXiv preprint arXiv:2603.27440},
  year   = {2026}
}

Comments

10 pages, 6 figures, 4 tables. Submitted to EDM 2026

R2 v1 2026-07-01T11:42:32.899Z