English

Analyzing Cognitive Plausibility of Subword Tokenization

Computation and Language 2023-10-23 v1

Abstract

Subword tokenization has become the de-facto standard for tokenization, although comparative evaluations of subword vocabulary quality across languages are scarce. Existing evaluation studies focus on the effect of a tokenization algorithm on the performance in downstream tasks, or on engineering criteria such as the compression rate. We present a new evaluation paradigm that focuses on the cognitive plausibility of subword tokenization. We analyze the correlation of the tokenizer output with the response time and accuracy of human performance on a lexical decision task. We compare three tokenization algorithms across several languages and vocabulary sizes. Our results indicate that the UnigramLM algorithm yields less cognitively plausible tokenization behavior and a worse coverage of derivational morphemes, in contrast with prior work.

Keywords

Cite

@article{arxiv.2310.13348,
  title  = {Analyzing Cognitive Plausibility of Subword Tokenization},
  author = {Lisa Beinborn and Yuval Pinter},
  journal= {arXiv preprint arXiv:2310.13348},
  year   = {2023}
}

Comments

EMNLP 2023 (main)

R2 v1 2026-06-28T12:56:37.481Z