BioProBench:生物协议理解与推理的全面数据集与基准
计算与语言
2026-01-22 v3
摘要
自主科学实验的实现目前受限于大语言模型(LLM)难以把握生物协议所需的严格程序逻辑和准确性。为此,我们提出了 **BioProBench**——生物程序推理的全面资源。BioProBench 以 **BioProCorpus** 为基础,这是一个包含 27,000 份人工编写协议的基础语料库。从该语料库中,我们系统性地构建了超过 55 万个任务实例的数据集,既提供大规模的训练资源,又具备新颖指标的严格基准测试。评估了 10 个主流大语言模型,我们发现虽然一般理解能力较高,但在需要深层推理、量化精确性和安全性 awareness 的任务上表现显著下降。为证明 BioProCorpus 在缓解这些问题方面的价值,我们 developed了 **ProAgent**,基于该语料库,ProAgent 显著推进了研究领域的最新进展。BioProBench 提供了严谨的诊断性基准测试,以及为开发下一代可靠科学人工智能提供的基础资源。代码和数据已提供:https://github.com/YuyangSunshine/bioprotocolbench 和 https://huggingface.co/datasets/BioProBench/BioProBench。
关键词
引用
@article{arxiv.2505.07889,
title = {BioProBench: Comprehensive Dataset and Benchmark in Biological Protocol Understanding and Reasoning},
author = {Yuyang Liu and Liuzhenghao Lv and Xiancheng Zhang and Jingya Wang Li Yuan and Yonghong Tian},
journal= {arXiv preprint arXiv:2505.07889},
year = {2026}
}