促进乌托-阿兹特克语技术发展:以濒危的 Comanche 语为案例研究
计算与语言
2025-05-27 v1 机器学习
摘要
endangered 语言的数字排除仍是 NLP 中的关键挑战,限制了语言研究和复兴努力。本研究介绍了对 Comanche——一门濒临灭绝的乌托-阿兹特克语——的首次计算研究,展示了如何通过最小成本、社区导向的 NLP 干预来支持语言保存。我们呈现了一个包含 412 个短语的人工精选数据集、一个合成数据生成管道,以及对 GPT-4o 和 GPT-4o-mini 用于语言识别的实证评估。我们的实验结果表明,尽管在零样本设置下 LLM 对 Comanche 表现不佳,但通过 few-shot 提示显著提高性能,只需提供五个示例即可实现接近完美的准确率。我们的发现凸显了在低资源环境中有针对性的 NLP 方法的潜力,强调可见性是走向包容的第一步。通过为 Comanche 在 NLP 领域奠定基础,我们倡导优先考虑可访问性、文化敏感性和社区参与的计算方法。
引用
@article{arxiv.2505.18159,
title = {Advancing Uto-Aztecan Language Technologies: A Case Study on the Endangered Comanche Language},
author = {Jesus Alvarez C and Daua D. Karajeanes and Ashley Celeste Prado and John Ruttan and Ivory Yang and Sean O'Brien and Vasu Sharma and Kevin Zhu},
journal= {arXiv preprint arXiv:2505.18159},
year = {2025}
}
备注
11 pages, 13 figures; published in Proceedings of the Fifth Workshop on NLP for Indigenous Languages of the Americas (AmericasNLP 2025) at NAACL 2025, Albuquerque, NM