English

Language Diversity: Visible to Humans, Exploitable by Machines

Computation and Language 2022-03-10 v1

Abstract

The Universal Knowledge Core (UKC) is a large multilingual lexical database with a focus on language diversity and covering over a thousand languages. The aim of the database, as well as its tools and data catalogue, is to make the somewhat abstract notion of diversity visually understandable for humans and formally exploitable by machines. The UKC website lets users explore millions of individual words and their meanings, but also phenomena of cross-lingual convergence and divergence, such as shared interlingual meanings, lexicon similarities, cognate clusters, or lexical gaps. The UKC LiveLanguage Catalogue, in turn, provides access to the underlying lexical data in a computer-processable form, ready to be reused in cross-lingual applications.

Keywords

Cite

@article{arxiv.2203.04723,
  title  = {Language Diversity: Visible to Humans, Exploitable by Machines},
  author = {Gábor Bella and Erdenebileg Byambadorj and Yamini Chandrashekar and Khuyagbaatar Batsuren and Danish Ashgar Cheema and Fausto Giunchiglia},
  journal= {arXiv preprint arXiv:2203.04723},
  year   = {2022}
}

Comments

Accepted for publication in ACL 2022

R2 v1 2026-06-24T10:07:19.244Z