English

BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages

Computation and Language 2017-10-09 v1

Abstract

We present BPEmb, a collection of pre-trained subword unit embeddings in 275 languages, based on Byte-Pair Encoding (BPE). In an evaluation using fine-grained entity typing as testbed, BPEmb performs competitively, and for some languages bet- ter than alternative subword approaches, while requiring vastly fewer resources and no tokenization. BPEmb is available at https://github.com/bheinzerling/bpemb

Keywords

Cite

@article{arxiv.1710.02187,
  title  = {BPEmb: Tokenization-free Pre-trained Subword Embeddings in 275 Languages},
  author = {Benjamin Heinzerling and Michael Strube},
  journal= {arXiv preprint arXiv:1710.02187},
  year   = {2017}
}
R2 v1 2026-06-22T22:05:07.825Z