中文

从公共语料库中挖掘大规模领域特定知识

计算与语言 2025-05-27 v4

摘要

大型语言模型(LLMs)在各种任务中展现出巨大潜力,然而,针对特定领域的开源模型和数据仍然严重缺乏。先前的工作主要集中在手动指定资源并为特定领域收集高质量数据,这极其耗时且费力。为了解决这一局限,我们将大型模型引入数据收集流程,以指导领域特定信息的生成,并从大型公共语料库 Common Crawl(CC)中检索相关数据。我们将这种方法称为 Retrieve-from-CC。它不仅收集与领域特定知识相关的数据,还从公共语料库中挖掘包含潜在推理过程的数据。通过应用此方法,我们收集了一个名为 Retrieve-Pile 的知识领域相关数据集,该数据集涵盖科学、人文和其他类别等四个主要领域。通过分析,Retrieve-from-CC 能够有效地从所涵盖的知识领域中检索相关数据,并显著提高数学和知识相关推理能力测试的性能。我们已在 https://huggingface.co/datasets/Query-of-CC/Retrieve-Pile 上发布了 Retrieve-Pile。

关键词

引用

@article{arxiv.2401.14624,
  title  = {Unearthing Large Scale Domain-Specific Knowledge from Public Corpora},
  author = {Zhaoye Fei and Yunfan Shao and Linyang Li and Zhiyuan Zeng and Conghui He and Qipeng Guo and Hang Yan and Dahua Lin and Xipeng Qiu},
  journal= {arXiv preprint arXiv:2401.14624},
  year   = {2025}
}

备注

We have released the full data (total of 735GB) in https://huggingface.co/datasets/Query-of-CC/Retrieve-Pile and partial data (about 40GB) in https://huggingface.co/datasets/Query-of-CC/knowledge_pile