English

What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification

Computation and Language 2025-10-07 v1

Abstract

Generative large language models (LLMs) are now used extensively for text classification in computational social science (CSS). In this work, focus on the steps before and after LLM prompting -- conceptualization of concepts to be classified and using LLM predictions in downstream statistical inference -- which we argue have been overlooked in much of LLM-era CSS. We claim LLMs can tempt analysts to skip the conceptualization step, creating conceptualization errors that bias downstream estimates. Using simulations, we show that this conceptualization-induced bias cannot be corrected for solely by increasing LLM accuracy or post-hoc bias correction methods. We conclude by reminding CSS analysts that conceptualization is still a first-order concern in the LLM-era and provide concrete advice on how to pursue low-cost, unbiased, low-variance downstream estimates.

Keywords

Cite

@article{arxiv.2510.03541,
  title  = {What is a protest anyway? Codebook conceptualization is still a first-order concern in LLM-era classification},
  author = {Andrew Halterman and Katherine A. Keith},
  journal= {arXiv preprint arXiv:2510.03541},
  year   = {2025}
}
R2 v1 2026-07-01T06:16:30.084Z