English

A meta-analysis on the performance of machine-learning based language models for sentiment analysis

Computation and Language 2025-09-15 v1 Machine Learning Applications

Abstract

This paper presents a meta-analysis evaluating ML performance in sentiment analysis for Twitter data. The study aims to estimate the average performance, assess heterogeneity between and within studies, and analyze how study characteristics influence model performance. Using PRISMA guidelines, we searched academic databases and selected 195 trials from 20 studies with 12 study features. Overall accuracy, the most reported performance metric, was analyzed using double arcsine transformation and a three-level random effects model. The average overall accuracy of the AIC-optimized model was 0.80 [0.76, 0.84]. This paper provides two key insights: 1) Overall accuracy is widely used but often misleading due to its sensitivity to class imbalance and the number of sentiment classes, highlighting the need for normalization. 2) Standardized reporting of model performance, including reporting confusion matrices for independent test sets, is essential for reliable comparisons of ML classifiers across studies, which seems far from common practice.

Keywords

Cite

@article{arxiv.2509.09728,
  title  = {A meta-analysis on the performance of machine-learning based language models for sentiment analysis},
  author = {Elena Rohde and Jonas Klingwort and Christian Borgs},
  journal= {arXiv preprint arXiv:2509.09728},
  year   = {2025}
}
R2 v1 2026-07-01T05:32:33.661Z