English

Universal Language Model Fine-Tuning with Subword Tokenization for Polish

Computation and Language 2018-10-25 v1 Machine Learning Machine Learning

Abstract

Universal Language Model for Fine-tuning [arXiv:1801.06146] (ULMFiT) is one of the first NLP methods for efficient inductive transfer learning. Unsupervised pretraining results in improvements on many NLP tasks for English. In this paper, we describe a new method that uses subword tokenization to adapt ULMFiT to languages with high inflection. Our approach results in a new state-of-the-art for the Polish language, taking first place in Task 3 of PolEval'18. After further training, our final model outperformed the second best model by 35%. We have open-sourced our pretrained models and code.

Keywords

Cite

@article{arxiv.1810.10222,
  title  = {Universal Language Model Fine-Tuning with Subword Tokenization for Polish},
  author = {Piotr Czapla and Jeremy Howard and Marcin Kardas},
  journal= {arXiv preprint arXiv:1810.10222},
  year   = {2018}
}

Comments

PolEval 2018 Workshop

R2 v1 2026-06-23T04:50:52.680Z