Evaluating KGR10 Polish word embeddings in the recognition of temporal expressions using BiLSTM-CRF
Computation and Language
2019-04-09 v1 Machine Learning
Machine Learning
Abstract
The article introduces a new set of Polish word embeddings, built using KGR10 corpus, which contains more than 4 billion words. These embeddings are evaluated in the problem of recognition of temporal expressions (timexes) for the Polish language. We described the process of KGR10 corpus creation and a new approach to the recognition problem using Bidirectional Long-Short Term Memory (BiLSTM) network with additional CRF layer, where specific embeddings are essential. We presented experiments and conclusions drawn from them.
Keywords
Cite
@article{arxiv.1904.04055,
title = {Evaluating KGR10 Polish word embeddings in the recognition of temporal expressions using BiLSTM-CRF},
author = {Jan Kocoń and Michał Gawor},
journal= {arXiv preprint arXiv:1904.04055},
year = {2019}
}
Comments
Presented at TFML 2019 (Theoretical Foundations of Machine Learning)