ORPHEAS:面向检索增强生成的跨语言希腊语-英语嵌入模型
计算与语言
2026-04-23 v1 人工智能
摘要
在希腊语-英语双语应用中实现有效的检索增强生成,需要能够捕捉特定领域语义关系和跨语言语义对齐的嵌入模型。现有的多语言嵌入模型将其表征能力分散于众多语言,限制了其对希腊语的优化,且无法编码希腊语文本中固有的形态复杂性和特定领域术语结构。本文提出 ORPHEAS,一个专用于双语检索增强生成的希腊语-英语嵌入模型。ORPHEAS 使用基于知识图谱的微调方法生成的高质量数据集进行训练,该方法应用于多样化的多领域语料库,从而实现了语言无关的语义表征。在单语和跨语言检索基准上的数值实验表明,ORPHEAS 优于最先进的多语言嵌入模型,证明了对形态复杂语言进行领域特化微调不会损害其跨语言检索能力。
引用
@article{arxiv.2604.20666,
title = {ORPHEAS: A Cross-Lingual Greek-English Embedding Model for Retrieval-Augmented Generation},
author = {Ioannis E. Livieris and Athanasios Koursaris and Alexandra Apostolopoulou and Konstantinos Kanaris Dimitris Tsakalidis and George Domalis},
journal= {arXiv preprint arXiv:2604.20666},
year = {2026}
}
备注
This paper has been accepted for presentation at Engineering Applications and Advances of Artificial Intelligence 2026 (EAAAI'26)