Towards Red Teaming in Multimodal and Multilingual Translation
Abstract
Assessing performance in Natural Language Processing is becoming increasingly complex. One particular challenge is the potential for evaluation datasets to overlap with training data, either directly or indirectly, which can lead to skewed results and overestimation of model performance. As a consequence, human evaluation is gaining increasing interest as a means to assess the performance and reliability of models. One such method is the red teaming approach, which aims to generate edge cases where a model will produce critical errors. While this methodology is becoming standard practice for generative AI, its application to the realm of conditional AI remains largely unexplored. This paper presents the first study on human-based red teaming for Machine Translation (MT), marking a significant step towards understanding and improving the performance of translation models. We delve into both human-based red teaming and a study on automation, reporting lessons learned and providing recommendations for both translation models and red teaming drills. This pioneering work opens up new avenues for research and development in the field of MT.
Cite
@article{arxiv.2401.16247,
title = {Towards Red Teaming in Multimodal and Multilingual Translation},
author = {Christophe Ropers and David Dale and Prangthip Hansanti and Gabriel Mejia Gonzalez and Ivan Evtimov and Corinne Wong and Christophe Touret and Kristina Pereyra and Seohyun Sonia Kim and Cristian Canton Ferrer and Pierre Andrews and Marta R. Costa-jussà},
journal= {arXiv preprint arXiv:2401.16247},
year = {2024}
}
Comments
arXiv admin note: substantial text overlap with arXiv:2312.05187