English

A Case Study of Scalable Content Annotation Using Multi-LLM Consensus and Human Review

Human-Computer Interaction 2025-04-29 v3

Abstract

Content annotation at scale remains challenging, requiring substantial human expertise and effort. This paper presents a case study in code documentation analysis, where we explore the balance between automation efficiency and annotation accuracy. We present MCHR (Multi-LLM Consensus with Human Review), a novel semi-automated framework that enhances annotation scalability through the systematic integration of multiple LLMs and targeted human review. Our framework introduces a structured consensus-building mechanism among LLMs and an adaptive review protocol that strategically engages human expertise. Through our case study, we demonstrate that MCHR reduces annotation time by 32% to 100% compared to manual annotation while maintaining high accuracy (85.5% to 98%) across different difficulty levels, from basic binary classification to challenging open-set scenarios.

Keywords

Cite

@article{arxiv.2503.17620,
  title  = {A Case Study of Scalable Content Annotation Using Multi-LLM Consensus and Human Review},
  author = {Mingyue Yuan and Jieshan Chen and Zhenchang Xing and Gelareh Mohammadi and Aaron Quigley},
  journal= {arXiv preprint arXiv:2503.17620},
  year   = {2025}
}

Comments

7 pages, GenAICHI: CHI 2025 Workshop on Generative AI and HCI

R2 v1 2026-06-28T22:30:38.185Z