English

LAST at SemEval-2021 Task 1: Improving Multi-Word Complexity Prediction Using Bigram Association Measures

Computation and Language 2021-05-21 v1

Abstract

This paper describes the system developed by the Laboratoire d'analyse statistique des textes (LAST) for the Lexical Complexity Prediction shared task at SemEval-2021. The proposed system is made up of a LightGBM model fed with features obtained from many word frequency lists, published lexical norms and psychometric data. For tackling the specificity of the multi-word task, it uses bigram association measures. Despite that the only contextual feature used was sentence length, the system achieved an honorable performance in the multi-word task, but poorer in the single word task. The bigram association measures were found useful, but to a limited extent.

Keywords

Cite

@article{arxiv.2105.09653,
  title  = {LAST at SemEval-2021 Task 1: Improving Multi-Word Complexity Prediction Using Bigram Association Measures},
  author = {Yves Bestgen},
  journal= {arXiv preprint arXiv:2105.09653},
  year   = {2021}
}

Comments

Accepted at SemEval-2021

R2 v1 2026-06-24T02:17:48.423Z