English

Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews

Computation and Language 2025-09-19 v1

Abstract

Understanding emotions in the Indonesian language is essential for improving customer experiences in e-commerce. This study focuses on enhancing the accuracy of emotion classification in Indonesian by leveraging advanced language models, IndoBERT and DistilBERT. A key component of our approach was data processing, specifically data augmentation, which included techniques such as back-translation and synonym replacement. These methods played a significant role in boosting the model's performance. After hyperparameter tuning, IndoBERT achieved an accuracy of 80\%, demonstrating the impact of careful data processing. While combining multiple IndoBERT models led to a slight improvement, it did not significantly enhance performance. Our findings indicate that IndoBERT was the most effective model for emotion classification in Indonesian, with data augmentation proving to be a vital factor in achieving high accuracy. Future research should focus on exploring alternative architectures and strategies to improve generalization for Indonesian NLP tasks.

Keywords

Cite

@article{arxiv.2509.14611,
  title  = {Leveraging IndoBERT and DistilBERT for Indonesian Emotion Classification in E-Commerce Reviews},
  author = {William Christian and Daniel Adamlu and Adrian Yu and Derwin Suhartono},
  journal= {arXiv preprint arXiv:2509.14611},
  year   = {2025}
}