English

Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate

Computation and Language 2025-06-05 v1 Artificial Intelligence Computers and Society Human-Computer Interaction Machine Learning

Abstract

Automated counter-narratives (CN) offer a promising strategy for mitigating online hate speech, yet concerns about their affective tone, accessibility, and ethical risks remain. We propose a framework for evaluating Large Language Model (LLM)-generated CNs across four dimensions: persona framing, verbosity and readability, affective tone, and ethical robustness. Using GPT-4o-Mini, Cohere's CommandR-7B, and Meta's LLaMA 3.1-70B, we assess three prompting strategies on the MT-Conan and HatEval datasets. Our findings reveal that LLM-generated CNs are often verbose and adapted for people with college-level literacy, limiting their accessibility. While emotionally guided prompts yield more empathetic and readable responses, there remain concerns surrounding safety and effectiveness.

Keywords

Cite

@article{arxiv.2506.04043,
  title  = {Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate},
  author = {Mikel K. Ngueajio and Flor Miriam Plaza-del-Arco and Yi-Ling Chung and Danda B. Rawat and Amanda Cercas Curry},
  journal= {arXiv preprint arXiv:2506.04043},
  year   = {2025}
}

Comments

Accepted at ACL WOAH 2025