中文

泰语OCR基准测试:面向泰语视觉语言理解的任务多样化基准

计算与语言 2025-12-05 v3

摘要

我们提出了 ThaiOCRBench,这是首个综合性基准,用于评估视觉语言模型(VLM)在泰语文本丰富的视觉理解任务上的表现。尽管多模态建模近年来取得了进展,但现有基准主要聚焦于高资源语言,泰语在需要文档结构理解的任务中仍严重不足。ThaiOCRBench 通过提供覆盖 13 个任务类别的 2,808 个人工标注样本,弥补了这一差距。我们在零样态设置下评估了广泛的前沿 VLM,包括专有和开源系统。结果显示,专有模型(如 Gemini 2.5 Pro)在性能上显著领先于开源模型。值得注意的是,细粒度文本识别和手写内容提取在开源模型中表现下降最大。通过详细的错误分析,我们识别出关键挑战包括语言偏见、结构不匹配和幻觉内容。ThaiOCRBench 提供了评估低资源、脚本复杂环境下 VLM 的标准化框架,并为改进泰语文档理解提供了可操作性建议。

关键词

引用

@article{arxiv.2511.04479,
  title  = {ThaiOCRBench: A Task-Diverse Benchmark for Vision-Language Understanding in Thai},
  author = {Surapon Nonesung and Teetouch Jaknamon and Sirinya Chaiophat and Natapong Nitarach and Chanakan Wittayasakpan and Warit Sirichotedumrong and Adisai Na-Thalang and Kunat Pipatanakul},
  journal= {arXiv preprint arXiv:2511.04479},
  year   = {2025}
}

备注

Accepted at IJCNLP-AACL 2025 (Main). This version includes the corrected Table 2 and an updated conclusion regarding the deletion count of the Gemma model