English

Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis

Computation and Language 2025-12-15 v1

Abstract

We investigated the impact of data selection on machine translation fine-tuning for open LLMs. Using Japanese-English corpora, we compare five selectors: TF-IDF, COMET Kiwi, QuRate, FD-Score, and random selection, under controlled training conditions. We observed that semantic selectors consistently outperform lexical and geometry-based heuristics, and that even when the selected data differ by less than 3%, the impact on model performance is substantial, underscoring the sensitivity of fine-tuning to data quality.

Keywords

Cite

@article{arxiv.2512.11388,
  title  = {Improving Translation Quality by Selecting Better Data for LLM Fine-Tuning: A Comparative Analysis},
  author = {Felipe Ribeiro Fujita de Mello and Hideyuki Takada},
  journal= {arXiv preprint arXiv:2512.11388},
  year   = {2025}
}

Comments

To appear at IEEE Big Data 2025

R2 v1 2026-07-01T08:21:57.938Z