中文

LLM 中的翻译不对称性作为数据增强因子:6 种罗曼什语变体的案例研究

计算与语言 2026-03-27 v1

摘要

低资源机器翻译的最新策略依赖 LLM 从高资源语言生成合成数据。我们发现该方法对罗曼什语失效,因为 LLM 倾向于混淆其 6 种不同语言变体。我们的实验表明,数据增强的方向应与源语言和目标语言之间的资源梯度对齐。该方法在罗曼什语资源最少的变体上超越了 Gemini 3 Pro 23 BLEU。人类评估确认,我们的实验产生了首个能在 individual 罗曼什语变体中生成流畅翻译的模型。

关键词

引用

@article{arxiv.2603.25489,
  title  = {Translation Asymmetry in LLMs as a Data Augmentation Factor: A Case Study for 6 Romansh Language Varieties},
  author = {Jannis Vamvas and Ignacio Pérez Prat and Angela Heldstab and Dominic P. Fischer and Sina Ahmadi and Rico Sennrich},
  journal= {arXiv preprint arXiv:2603.25489},
  year   = {2026}
}

备注

Preprint