English

EnTaCs: Analyzing the Relationship Between Sentiment and Language Choice in English-Tamil Code-Switching

Computation and Language 2026-03-30 v1

Abstract

This paper investigates the relationship between utterance sentiment and language choice in English-Tamil code-switched text, using methods from machine learning and statistical modelling. We apply a fine-tuned XLM-RoBERTa model for token-level language identification on 35,650 romanized YouTube comments from the DravidianCodeMix dataset, producing per-utterance measurements of English proportion and language switch frequency. Linear regression analysis reveals that positive utterances exhibit significantly greater English proportion (34.3%) than negative utterances (24.8%), and mixed-sentiment utterances show the highest language switch frequency when controlling for utterance length. These findings support the hypothesis that emotional content demonstrably influences language choice in multilingual code-switching settings, due to socio-linguistic associations of prestige and identity with embedded and matrix languages.

Keywords

Cite

@article{arxiv.2603.26587,
  title  = {EnTaCs: Analyzing the Relationship Between Sentiment and Language Choice in English-Tamil Code-Switching},
  author = {Paul Bontempo},
  journal= {arXiv preprint arXiv:2603.26587},
  year   = {2026}
}

Comments

5 pages, 2 figures

R2 v1 2026-07-01T11:41:06.802Z