English

Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora

Computation and Language 2016-09-27 v2

Abstract

A word's sentiment depends on the domain in which it is used. Computational social science research thus requires sentiment lexicons that are specific to the domains being studied. We combine domain-specific word embeddings with a label propagation framework to induce accurate domain-specific sentiment lexicons using small sets of seed words, achieving state-of-the-art performance competitive with approaches that rely on hand-curated resources. Using our framework we perform two large-scale empirical studies to quantify the extent to which sentiment varies across time and between communities. We induce and release historical sentiment lexicons for 150 years of English and community-specific sentiment lexicons for 250 online communities from the social media forum Reddit. The historical lexicons show that more than 5% of sentiment-bearing (non-neutral) English words completely switched polarity during the last 150 years, and the community-specific lexicons highlight how sentiment varies drastically between different communities.

Keywords

Cite

@article{arxiv.1606.02820,
  title  = {Inducing Domain-Specific Sentiment Lexicons from Unlabeled Corpora},
  author = {William L. Hamilton and Kevin Clark and Jure Leskovec and Dan Jurafsky},
  journal= {arXiv preprint arXiv:1606.02820},
  year   = {2016}
}

Comments

11 pages, 5 figures, EMNLP 2016

R2 v1 2026-06-22T14:21:22.064Z