中文

评估大型语言模型中的文化知识处理:集成检索增强生成的认知基准框架

计算与语言 2025-11-04 v1

摘要

本研究提出了一种认知基准框架,用于评估大型语言模型 (LLM) 处理和应用特定文化知识的方式。该框架将霍布斯分类法与检索增强生成 (RAG) 集成,用于衡量模型在六个层次认知领域:记忆 (Remembering)、理解 (Understanding)、应用 (Applying)、分析 (Analyzing)、评估 (Evaluating) 和创造 (Creating) 中的性能。该评估以精选的台湾客家数字文化档案库为主要测试基准,衡量 LLM 生成的响应在语义准确性和文化相关性方面的表现。

关键词

引用

@article{arxiv.2511.01649,
  title  = {Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation},
  author = {Hung-Shin Lee and Chen-Chi Chang and Ching-Yuan Chen and Yun-Hsiang Hsu},
  journal= {arXiv preprint arXiv:2511.01649},
  year   = {2025}
}

备注

This paper has been accepted by The Electronic Library, and the full article is now available on Emerald Insight