English

Evaluating CxG Generalisation in LLMs via Construction-Based NLI Fine Tuning

Computation and Language 2025-09-23 v1

Abstract

We probe large language models' ability to learn deep form-meaning mappings as defined by construction grammars. We introduce the ConTest-NLI benchmark of 80k sentences covering eight English constructions from highly lexicalized to highly schematic. Our pipeline generates diverse synthetic NLI triples via templating and the application of a model-in-the-loop filter. This provides aspects of human validation to ensure challenge and label reliability. Zero-shot tests on leading LLMs reveal a 24% drop in accuracy between naturalistic (88%) and adversarial data (64%), with schematic patterns proving hardest. Fine-tuning on a subset of ConTest-NLI yields up to 9% improvement, yet our results highlight persistent abstraction gaps in current LLMs and offer a scalable framework for evaluating construction-informed learning.

Keywords

Cite

@article{arxiv.2509.16422,
  title  = {Evaluating CxG Generalisation in LLMs via Construction-Based NLI Fine Tuning},
  author = {Tom Mackintosh and Harish Tayyar Madabushi and Claire Bonial},
  journal= {arXiv preprint arXiv:2509.16422},
  year   = {2025}
}
R2 v1 2026-07-01T05:46:41.878Z