English

PerQ: Efficient Evaluation of Multilingual Text Personalization Quality

Computation and Language 2025-10-01 v1 Artificial Intelligence

Abstract

Since no metrics are available to evaluate specific aspects of a text, such as its personalization quality, the researchers often rely solely on large language models to meta-evaluate such texts. Due to internal biases of individual language models, it is recommended to use multiple of them for combined evaluation, which directly increases costs of such meta-evaluation. In this paper, a computationally efficient method for evaluation of personalization quality of a given text (generated by a language model) is introduced, called PerQ. A case study of comparison of generation capabilities of large and small language models shows the usability of the proposed metric in research, effectively reducing the waste of resources.

Keywords

Cite

@article{arxiv.2509.25903,
  title  = {PerQ: Efficient Evaluation of Multilingual Text Personalization Quality},
  author = {Dominik Macko and Andrew Pulver},
  journal= {arXiv preprint arXiv:2509.25903},
  year   = {2025}
}
R2 v1 2026-07-01T06:07:01.791Z