English

Towards Effective Ancient Chinese Translation: Dataset, Model, and Evaluation

Computation and Language 2023-08-02 v1

Abstract

Interpreting ancient Chinese has been the key to comprehending vast Chinese literature, tradition, and civilization. In this paper, we propose Erya for ancient Chinese translation. From a dataset perspective, we collect, clean, and classify ancient Chinese materials from various sources, forming the most extensive ancient Chinese resource to date. From a model perspective, we devise Erya training method oriented towards ancient Chinese. We design two jointly-working tasks: disyllabic aligned substitution (DAS) and dual masked language model (DMLM). From an evaluation perspective, we build a benchmark to judge ancient Chinese translation quality in different scenarios and evaluate the ancient Chinese translation capacities of various existing models. Our model exhibits remarkable zero-shot performance across five domains, with over +12.0 BLEU against GPT-3.5 models and better human evaluation results than ERNIE Bot. Subsequent fine-tuning further shows the superior transfer capability of Erya model with +6.2 BLEU gain. We release all the above-mentioned resources at https://github.com/RUCAIBox/Erya.

Keywords

Cite

@article{arxiv.2308.00240,
  title  = {Towards Effective Ancient Chinese Translation: Dataset, Model, and Evaluation},
  author = {Geyang Guo and Jiarong Yang and Fengyuan Lu and Jiaxin Qin and Tianyi Tang and Wayne Xin Zhao},
  journal= {arXiv preprint arXiv:2308.00240},
  year   = {2023}
}

Comments

Accepted by NLPCC 2023

R2 v1 2026-06-28T11:45:07.487Z