基于 BERT 的越南语言事实验证数据集模型
计算与语言
2025-03-04 v1 人工智能
摘要
信息和通信技术的快速发展已便利了信息的获取。然而,这一进步也 necessitate 更严格的验证措施来确保信息的准确性,尤其是在越南语境下。本文提出了一种方法,旨在解决针对越南语言事实验证数据集的挑战,通过将句子选择模块和分类模块集成到统一的网络架构中。该方法利用大语言模型的能力,采用预训练的 PhoBERT 和 XLM-RoBERTa 作为网络的 backbone。该模型在名为 ISE-DSC01 的越南语言数据集上进行训练,相较于基线模型在所有三个指标上均表现出优异性能。值得注意的是,我们实现了 75.11% 的 Strict Accuracy 水平,相当于基线模型提升了 28.83%。
引用
@article{arxiv.2503.00356,
title = {BERT-based model for Vietnamese Fact Verification Dataset},
author = {Bao Tran and T. N. Khanh and Khang Nguyen Tuong and Thien Dang and Quang Nguyen and Nguyen T. Thinh and Vo T. Hung},
journal= {arXiv preprint arXiv:2503.00356},
year = {2025}
}
备注
accepted for Oral Presentation in CITA 2024 (The 13th Conference on Information Technology and Its Applications) and will be published in VOLUME 1 OF CITA 2024 (Volume of the Lecture Notes in Network and Systems, Springer)