大语言模型分词器在阿萨姆语中的性能评估
计算与语言
2025-04-08 v1
摘要
分词器的训练在深度学习模型的性能中发挥重要作用。本研究旨在了解五个最先进(SOTA)大语言模型(LLM)在印度阿萨姆语中的分词器性能。该研究对于理解来自低资源语言(如阿萨姆语)的多语言支持至关重要。我们的研究结果表明,来自 Two AI 的 SUTRA 的分词器性能最佳,平均归一化序列长度(NSL)值为 0.45,其后是来自 Open AI 的 GPT-4o 的分词器,平均 NSL 值为 0.54,其后是 Gemma 2、Meta Llama 3.1 和 Mistral Large Instruct 2407,分别具有平均 NSL 值为 0.82、1.4 和 1.48。
引用
@article{arxiv.2410.03718,
title = {Performance Evaluation of Tokenizers in Large Language Models for the Assamese Language},
author = {Sagar Tamang and Dibya Jyoti Bora},
journal= {arXiv preprint arXiv:2410.03718},
year = {2025}
}