English

Towards Large Language Model driven Reference-less Translation Evaluation for English and Indian Languages

Computation and Language 2024-04-04 v1

Abstract

With the primary focus on evaluating the effectiveness of large language models for automatic reference-less translation assessment, this work presents our experiments on mimicking human direct assessment to evaluate the quality of translations in English and Indian languages. We constructed a translation evaluation task where we performed zero-shot learning, in-context example-driven learning, and fine-tuning of large language models to provide a score out of 100, where 100 represents a perfect translation and 1 represents a poor translation. We compared the performance of our trained systems with existing methods such as COMET, BERT-Scorer, and LABSE, and found that the LLM-based evaluator (LLaMA-2-13B) achieves a comparable or higher overall correlation with human judgments for the considered Indian language pairs.

Keywords

Cite

@article{arxiv.2404.02512,
  title  = {Towards Large Language Model driven Reference-less Translation Evaluation for English and Indian Languages},
  author = {Vandan Mujadia and Pruthwik Mishra and Arafat Ahsan and Dipti Misra Sharma},
  journal= {arXiv preprint arXiv:2404.02512},
  year   = {2024}
}

Comments

arXiv admin note: text overlap with arXiv:2311.09216

R2 v1 2026-06-28T15:42:41.870Z