评估大型语言模型中的文化知识处理:集成检索增强生成的认知基准框架
计算与语言
2025-11-04 v1
摘要
本研究提出了一种认知基准框架,用于评估大型语言模型 (LLM) 处理和应用特定文化知识的方式。该框架将霍布斯分类法与检索增强生成 (RAG) 集成,用于衡量模型在六个层次认知领域:记忆 (Remembering)、理解 (Understanding)、应用 (Applying)、分析 (Analyzing)、评估 (Evaluating) 和创造 (Creating) 中的性能。该评估以精选的台湾客家数字文化档案库为主要测试基准,衡量 LLM 生成的响应在语义准确性和文化相关性方面的表现。
引用
@article{arxiv.2511.01649,
title = {Evaluating Cultural Knowledge Processing in Large Language Models: A Cognitive Benchmarking Framework Integrating Retrieval-Augmented Generation},
author = {Hung-Shin Lee and Chen-Chi Chang and Ching-Yuan Chen and Yun-Hsiang Hsu},
journal= {arXiv preprint arXiv:2511.01649},
year = {2025}
}
备注
This paper has been accepted by The Electronic Library, and the full article is now available on Emerald Insight