English

Examination and Extension of Strategies for Improving Personalized Language Modeling via Interpolation

Computation and Language 2020-06-11 v1 Machine Learning

Abstract

In this paper, we detail novel strategies for interpolating personalized language models and methods to handle out-of-vocabulary (OOV) tokens to improve personalized language models. Using publicly available data from Reddit, we demonstrate improvements in offline metrics at the user level by interpolating a global LSTM-based authoring model with a user-personalized n-gram model. By optimizing this approach with a back-off to uniform OOV penalty and the interpolation coefficient, we observe that over 80% of users receive a lift in perplexity, with an average of 5.2% in perplexity lift per user. In doing this research we extend previous work in building NLIs and improve the robustness of metrics for downstream tasks.

Keywords

Cite

@article{arxiv.2006.05469,
  title  = {Examination and Extension of Strategies for Improving Personalized Language Modeling via Interpolation},
  author = {Liqun Shao and Sahitya Mantravadi and Tom Manzini and Alejandro Buendia and Manon Knoertzer and Soundar Srinivasan and Chris Quirk},
  journal= {arXiv preprint arXiv:2006.05469},
  year   = {2020}
}

Comments

ACL Natural Language Interface Workshop 2020, short paper