English

Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy

Computation and Language 2026-01-29 v1 Artificial Intelligence

Abstract

Open-ended question answering (QA) evaluates a model's ability to perform contextualized reasoning beyond factual recall. This challenge is especially acute in practice-based domains, where knowledge is procedural and grounded in professional judgment, while most existing LLM benchmarks depend on pre-existing human exam datasets that are often unavailable in such settings. We introduce a framework for automated benchmark generation from expert-authored guidelines informed by Bloom's Taxonomy. It converts expert practices into implicit violation-based scenarios and expands them into auto-graded multiple-choice questions (MCQs) and multi-turn dialogues across four cognitive levels, enabling deterministic, reproducible, and scalable evaluation. Applied to three applied domains: teaching, dietetics, and caregiving, we find differences between model and human-like reasoning: LLMs sometimes perform relatively better on higher-order reasoning (Analyze) but fail more frequently on lower-level items (Remember). We produce large-scale, psychometrically informed benchmarks that surface these non-intuitive model behaviors and enable evaluation of contextualized reasoning in real-world settings.

Keywords

Cite

@article{arxiv.2601.20253,
  title  = {Automated Benchmark Generation from Domain Guidelines Informed by Bloom's Taxonomy},
  author = {Si Chen and Le Huy Khiem and Annalisa Szymanski and Ronald Metoyer and Ting Hua and Nitesh V. Chawla},
  journal= {arXiv preprint arXiv:2601.20253},
  year   = {2026}
}
R2 v1 2026-07-01T09:23:15.892Z