Sentiment Analysis of Code-Mixed Languages leveraging Resource Rich Languages
Abstract
Code-mixed data is an important challenge of natural language processing because its characteristics completely vary from the traditional structures of standard languages. In this paper, we propose a novel approach called Sentiment Analysis of Code-Mixed Text (SACMT) to classify sentences into their corresponding sentiment - positive, negative or neutral, using contrastive learning. We utilize the shared parameters of siamese networks to map the sentences of code-mixed and standard languages to a common sentiment space. Also, we introduce a basic clustering based preprocessing method to capture variations of code-mixed transliterated words. Our experiments reveal that SACMT outperforms the state-of-the-art approaches in sentiment analysis for code-mixed text by 7.6% in accuracy and 10.1% in F-score.
Cite
@article{arxiv.1804.00806,
title = {Sentiment Analysis of Code-Mixed Languages leveraging Resource Rich Languages},
author = {Nurendra Choudhary and Rajat Singh and Ishita Bindlish and Manish Shrivastava},
journal= {arXiv preprint arXiv:1804.00806},
year = {2024}
}
Comments
Accepted Long Paper at 19th International Conference on Computational Linguistics and Intelligent Text Processing, March 2018, Hanoi, Vietnam. arXiv admin note: text overlap with arXiv:1804.00805