English

CORI: CJKV Benchmark with Romanization Integration -- A step towards Cross-lingual Transfer Beyond Textual Scripts

Computation and Language 2024-04-22 v1 Artificial Intelligence Machine Learning

Abstract

Naively assuming English as a source language may hinder cross-lingual transfer for many languages by failing to consider the importance of language contact. Some languages are more well-connected than others, and target languages can benefit from transferring from closely related languages; for many languages, the set of closely related languages does not include English. In this work, we study the impact of source language for cross-lingual transfer, demonstrating the importance of selecting source languages that have high contact with the target language. We also construct a novel benchmark dataset for close contact Chinese-Japanese-Korean-Vietnamese (CJKV) languages to further encourage in-depth studies of language contact. To comprehensively capture contact between these languages, we propose to integrate Romanized transcription beyond textual scripts via Contrastive Learning objectives, leading to enhanced cross-lingual representations and effective zero-shot cross-lingual transfer.

Cite

@article{arxiv.2404.12618,
  title  = {CORI: CJKV Benchmark with Romanization Integration -- A step towards Cross-lingual Transfer Beyond Textual Scripts},
  author = {Hoang H. Nguyen and Chenwei Zhang and Ye Liu and Natalie Parde and Eugene Rohrbaugh and Philip S. Yu},
  journal= {arXiv preprint arXiv:2404.12618},
  year   = {2024}
}

Comments

Accepted at LREC-COLING 2024

R2 v1 2026-06-28T15:59:25.099Z