English

On the Learnability of Concepts: With Applications to Comparing Word Embedding Algorithms

Computation and Language 2020-06-18 v1 Artificial Intelligence Machine Learning

Abstract

Word Embeddings are used widely in multiple Natural Language Processing (NLP) applications. They are coordinates associated with each word in a dictionary, inferred from statistical properties of these words in a large corpus. In this paper we introduce the notion of "concept" as a list of words that have shared semantic content. We use this notion to analyse the learnability of certain concepts, defined as the capability of a classifier to recognise unseen members of a concept after training on a random subset of it. We first use this method to measure the learnability of concepts on pretrained word embeddings. We then develop a statistical analysis of concept learnability, based on hypothesis testing and ROC curves, in order to compare the relative merits of various embedding algorithms using a fixed corpora and hyper parameters. We find that all embedding methods capture the semantic content of those word lists, but fastText performs better than the others.

Keywords

Cite

@article{arxiv.2006.09896,
  title  = {On the Learnability of Concepts: With Applications to Comparing Word Embedding Algorithms},
  author = {Adam Sutton and Nello Cristianini},
  journal= {arXiv preprint arXiv:2006.09896},
  year   = {2020}
}

Comments

7 Pages. AIAI 2020. 5 equations 6 tables

R2 v1 2026-06-23T16:24:20.924Z