VenusFactory:蛋白工程数据检索与语言模型微调的统一平台
计算与语言
2025-03-20 v1 人工智能
定量方法
摘要
然然语言处理(NLP)已显著影响人类语言之外的科学领域,包括蛋白工程,其中预训练蛋白语言模型(PLMs)已显示出显著的成功。然而,跨学科的采用仍受到数据收集、任务基准测试和应用等挑战的限制。本工作提出了VenusFactory,一个多功能引擎,集成了生物学数据检索、标准化任务基准测试和模块化PLM微调。VenusFactory支持计算机科学和生物学社区,提供命令行执行和基于Gradio的无代码界面两种方式,集成40多个蛋白相关数据集和40多个流行PLMs。所有实现均开源于 https://github.com/tyang816/VenusFactory。
引用
@article{arxiv.2503.15438,
title = {VenusFactory: A Unified Platform for Protein Engineering Data Retrieval and Language Model Fine-Tuning},
author = {Yang Tan and Chen Liu and Jingyuan Gao and Banghao Wu and Mingchen Li and Ruilin Wang and Lingrong Zhang and Huiqun Yu and Guisheng Fan and Liang Hong and Bingxin Zhou},
journal= {arXiv preprint arXiv:2503.15438},
year = {2025}
}
备注
12 pages, 1 figure, 8 tables