English

Adaptive Ensembling: Unsupervised Domain Adaptation for Political Document Analysis

Computation and Language 2019-10-29 v1

Abstract

Insightful findings in political science often require researchers to analyze documents of a certain subject or type, yet these documents are usually contained in large corpora that do not distinguish between pertinent and non-pertinent documents. In contrast, we can find corpora that label relevant documents but have limitations (e.g., from a single source or era), preventing their use for political science research. To bridge this gap, we present \textit{adaptive ensembling}, an unsupervised domain adaptation framework, equipped with a novel text classification model and time-aware training to ensure our methods work well with diachronic corpora. Experiments on an expert-annotated dataset show that our framework outperforms strong benchmarks. Further analysis indicates that our methods are more stable, learn better representations, and extract cleaner corpora for fine-grained analysis.

Keywords

Cite

@article{arxiv.1910.12698,
  title  = {Adaptive Ensembling: Unsupervised Domain Adaptation for Political Document Analysis},
  author = {Shrey Desai and Barea Sinno and Alex Rosenfeld and Junyi Jessy Li},
  journal= {arXiv preprint arXiv:1910.12698},
  year   = {2019}
}

Comments

Accepted to EMNLP 2019

R2 v1 2026-06-23T11:57:12.569Z