English

An Empirical Study on Sentiment Classification of Chinese Review using Word Embedding

Computation and Language 2015-11-06 v1

Abstract

In this article, how word embeddings can be used as features in Chinese sentiment classification is presented. Firstly, a Chinese opinion corpus is built with a million comments from hotel review websites. Then the word embeddings which represent each comment are used as input in different machine learning methods for sentiment classification, including SVM, Logistic Regression, Convolutional Neural Network (CNN) and ensemble methods. These methods get better performance compared with N-gram models using Naive Bayes (NB) and Maximum Entropy (ME). Finally, a combination of machine learning methods is proposed which presents an outstanding performance in precision, recall and F1 score. After selecting the most useful methods to construct the combinational model and testing over the corpus, the final F1 score is 0.920.

Keywords

Cite

@article{arxiv.1511.01665,
  title  = {An Empirical Study on Sentiment Classification of Chinese Review using Word Embedding},
  author = {Yiou Lin and Hang Lei and Jia Wu and Xiaoyu Li},
  journal= {arXiv preprint arXiv:1511.01665},
  year   = {2015}
}

Comments

The 29th Pacific Asia Conference on Language, Information and Computing