English

Towards Red Teaming in Multimodal and Multilingual Translation

Computation and Language 2024-01-30 v1 Computers and Society

Abstract

Assessing performance in Natural Language Processing is becoming increasingly complex. One particular challenge is the potential for evaluation datasets to overlap with training data, either directly or indirectly, which can lead to skewed results and overestimation of model performance. As a consequence, human evaluation is gaining increasing interest as a means to assess the performance and reliability of models. One such method is the red teaming approach, which aims to generate edge cases where a model will produce critical errors. While this methodology is becoming standard practice for generative AI, its application to the realm of conditional AI remains largely unexplored. This paper presents the first study on human-based red teaming for Machine Translation (MT), marking a significant step towards understanding and improving the performance of translation models. We delve into both human-based red teaming and a study on automation, reporting lessons learned and providing recommendations for both translation models and red teaming drills. This pioneering work opens up new avenues for research and development in the field of MT.

Keywords

Cite

@article{arxiv.2401.16247,
  title  = {Towards Red Teaming in Multimodal and Multilingual Translation},
  author = {Christophe Ropers and David Dale and Prangthip Hansanti and Gabriel Mejia Gonzalez and Ivan Evtimov and Corinne Wong and Christophe Touret and Kristina Pereyra and Seohyun Sonia Kim and Cristian Canton Ferrer and Pierre Andrews and Marta R. Costa-jussà},
  journal= {arXiv preprint arXiv:2401.16247},
  year   = {2024}
}

Comments

arXiv admin note: substantial text overlap with arXiv:2312.05187

R2 v1 2026-06-28T14:30:22.076Z