English

The Imperfective Paradox in Large Language Models

Computation and Language 2026-04-23 v2

Abstract

Do Large Language Models (LLMs) genuinely grasp the compositional semantics of events, or do they rely on surface-level probabilistic heuristics? We investigate the Imperfective Paradox, a logical phenomenon where the past progressive aspect entails event realization for activities (e.g., running \to ran) but not for accomplishments (e.g., building \nrightarrow built). We introduce ImperfectiveNLI, a diagnostic dataset designed to probe this distinction across diverse semantic classes. Evaluating state-of-the-art open-weight models, we uncover a pervasive Teleological Bias: models systematically hallucinate completion for goal-oriented events, even overriding explicit textual cancellation. Prompting interventions partially reduce this bias but trigger a calibration crisis, causing models to incorrectly reject valid entailments for atelic verbs. Representational analyses further show that while internal embeddings often distinguish progressive from simple past forms, inference decisions are dominated by strong priors about goal attainment. Taken together, our findings indicate that these current open-weight LLMs operate as predictive narrative engines rather than faithful logical reasoners, and that resolving aspectual inference requires moving beyond prompting toward structurally grounded alignment.

Keywords

Cite

@article{arxiv.2601.09373,
  title  = {The Imperfective Paradox in Large Language Models},
  author = {Bolei Ma and Yusuke Miyao},
  journal= {arXiv preprint arXiv:2601.09373},
  year   = {2026}
}

Comments

ACL 2026

R2 v1 2026-07-01T09:04:09.387Z