English

Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence

Computation and Language 2025-07-15 v3

Abstract

This study investigates the linguistic understanding of Large Language Models (LLMs) regarding signifier (form) and signified (meaning) by distinguishing two LLM assessment paradigms: psycholinguistic and neurolinguistic. Traditional psycholinguistic evaluations often reflect statistical rules that may not accurately represent LLMs' true linguistic competence. We introduce a neurolinguistic approach, utilizing a novel method that combines minimal pair and diagnostic probing to analyze activation patterns across model layers. This method allows for a detailed examination of how LLMs represent form and meaning, and whether these representations are consistent across languages. We found: (1) Psycholinguistic and neurolinguistic methods reveal that language performance and competence are distinct; (2) Direct probability measurement may not accurately assess linguistic competence; (3) Instruction tuning won't change much competence but improve performance; (4) LLMs exhibit higher competence and performance in form compared to meaning. Additionally, we introduce new conceptual minimal pair datasets for Chinese (COMPS-ZH) and German (COMPS-DE), complementing existing English datasets.

Keywords

Cite

@article{arxiv.2411.07533,
  title  = {Large Language Models as Neurolinguistic Subjects: Discrepancy between Performance and Competence},
  author = {Linyang He and Ercong Nie and Helmut Schmid and Hinrich Schütze and Nima Mesgarani and Jonathan Brennan},
  journal= {arXiv preprint arXiv:2411.07533},
  year   = {2025}
}
R2 v1 2026-06-28T19:56:29.262Z