English

An English-Hindi Code-Mixed Corpus: Stance Annotation and Baseline System

Computation and Language 2018-05-31 v1

Abstract

Social media has become one of the main channels for peo- ple to communicate and share their views with the society. We can often detect from these views whether the person is in favor, against or neu- tral towards a given topic. These opinions from social media are very useful for various companies. We present a new dataset that consists of 3545 English-Hindi code-mixed tweets with opinion towards Demoneti- sation that was implemented in India in 2016 which was followed by a large countrywide debate. We present a baseline supervised classification system for stance detection developed using the same dataset that uses various machine learning techniques to achieve an accuracy of 58.7% on 10-fold cross validation.

Keywords

Cite

@article{arxiv.1805.11868,
  title  = {An English-Hindi Code-Mixed Corpus: Stance Annotation and Baseline System},
  author = {Sahil Swami and Ankush Khandelwal and Vinay Singh and Syed Sarfaraz Akhtar and Manish Shrivastava},
  journal= {arXiv preprint arXiv:1805.11868},
  year   = {2018}
}

Comments

9 pages, CICling 2018

R2 v1 2026-06-23T02:13:02.178Z