English

Your Students Don't Use LLMs Like You Wish They Did

Computation and Language 2026-04-28 v1 Computers and Society Human-Computer Interaction

Abstract

Educational NLP systems are typically evaluated using engagement metrics and satisfaction surveys, which are at best a proxy for meeting pedagogical goals. We introduce six computational metrics for automated evaluation of pedagogical alignment in student-AI dialogue. We validate our metrics through analysis of 12,650 messages across 500 conversations from four courses. Using our metrics, we identify a fundamental misalignment: educators design conversational tutors for sustained learning dialogue, but students mainly use them for answer-extraction. Deployment context is the strongest predictor of usage patterns, outweighing student preference or system design: when AI tools are optional, usage concentrates around deadlines; when integrated into course structure, students ask for solutions to verbatim assignment questions. Whole-dialogue evaluation misses these turn-by-turn patterns. Our metrics will enable researchers building educational dialogue systems to measure whether they are achieving their pedagogical goals.

Keywords

Cite

@article{arxiv.2604.23486,
  title  = {Your Students Don't Use LLMs Like You Wish They Did},
  author = {Sebastian Kobler and Matthew Clemson and Angela Sun and Jonathan K. Kummerfeld},
  journal= {arXiv preprint arXiv:2604.23486},
  year   = {2026}
}

Comments

To appear at ACL 2026 (Main Conference)