中文

DARWIN 1.5:作为材料科学自适应学习器的大语言模型

计算与语言 2025-05-22 v3

摘要

材料发现与设计旨在高度复杂且多样的物理空间中寻找具有理想性质的组成与结构。传统解决方案(如高通量模拟或机器学习)通常依赖于复杂的描述符,这阻碍了其在不同材料系统间的泛化性与可迁移性。此外,这些描述符可能无法充分表示宏观尺度的材料性质,而真实样本中的这些性质受结构缺陷和成分变化的影响,从而限制了其实际适用性。为了应对这些挑战,我们提出了 DARWIN 1.5,这是专为材料科学定制的最大开源大语言模型(LLM)。通过利用自然语言作为输入,DARWIN 消除了对特定任务描述符的需求,并为材料性质预测与发现提供了一种灵活、统一的方法。我们的方法整合了跨模态的 6M 篇材料领域论文和来自 49,256 种材料的 21 个实验数据集,同时实现了跨任务知识迁移。增强后的模型在基础 LLaMA-7B 架构上的预测准确率最高提升了 59.1%,并在 8 项材料设计任务中超越了 SOTA 机器学习方法。这些结果确立了 LLM 作为材料科学中开发通用且可扩展模型的有力基础。

关键词

引用

@article{arxiv.2412.11970,
  title  = {DARWIN 1.5: Large Language Models as Materials Science Adapted Learners},
  author = {Tong Xie and Yuwei Wan and Yixuan Liu and Yuchen Zeng and Shaozhou Wang and Wenjie Zhang and Clara Grazian and Chunyu Kit and Wanli Ouyang and Dongzhan Zhou and Bram Hoex},
  journal= {arXiv preprint arXiv:2412.11970},
  year   = {2025}
}

备注

This version of the manuscript was posted prematurely and contains inaccuracies that could mislead readers. The authors are preparing a significantly revised version with substantial methodological and experimental updates, and prefer to avoid confusion with earlier postings. We apologize for any inconvenience and thank the community for their understanding