English

Multilingual Models for Check-Worthy Social Media Posts Detection

Computation and Language 2024-08-14 v1

Abstract

This work presents an extensive study of transformer-based NLP models for detection of social media posts that contain verifiable factual claims and harmful claims. The study covers various activities, including dataset collection, dataset pre-processing, architecture selection, setup of settings, model training (fine-tuning), model testing, and implementation. The study includes a comprehensive analysis of different models, with a special focus on multilingual models where the same model is capable of processing social media posts in both English and in low-resource languages such as Arabic, Bulgarian, Dutch, Polish, Czech, Slovak. The results obtained from the study were validated against state-of-the-art models, and the comparison demonstrated the robustness of the proposed models. The novelty of this work lies in the development of multi-label multilingual classification models that can simultaneously detect harmful posts and posts that contain verifiable factual claims in an efficient way.

Keywords

Cite

@article{arxiv.2408.06737,
  title  = {Multilingual Models for Check-Worthy Social Media Posts Detection},
  author = {Sebastian Kula and Michal Gregor},
  journal= {arXiv preprint arXiv:2408.06737},
  year   = {2024}
}
R2 v1 2026-06-28T18:11:29.018Z