LLM 世界模型的专家评估:一种高 $T_c$ 超导体案例研究
摘要
大型语言模型(LLMs)显示出作为科学文献探索工具的巨大潜力。然而,其在提供针对特定领域中复杂问题的科学准确且全面答案方面的效果仍是活跃的研究领域。我们以高温铜系材料为例,对 LLM 系统的能力进行评估,以专家水平理解其文献。我们构建了一个覆盖该领域历史的 1,726 篇科学论文的数据库,并制定了 67 个专家提出的问题,以探测对文献的深层理解。我们 then evaluate six different LLM-based systems for answering these questions, including both commercial available closed models and a custom retrieval-augmented generation (RAG) system capable of retrieving images alongside text。 Experts then evaluate the answers of these systems against a rubric that assesses balanced perspectives, factual comprehensiveness, succinctness, and evidentiary support。 Among the six systems two using RAG on curated literature outperformed existing closed models across key metrics, particularly in providing comprehensive and well-supported answers。 We discuss promising aspects of LLM performances as well as critical short-comings of all the models。 The set of expert-formulated questions and the rubric will be valuable for assessing expert level performance of LLM based reasoning systems。
引用
@article{arxiv.2511.03782,
title = {Expert Evaluation of LLM World Models: A High-$T_c$ Superconductivity Case Study},
author = {Haoyu Guo and Maria Tikhanovskaya and Paul Raccuglia and Alexey Vlaskin and Chris Co and Daniel J. Liebling and Scott Ellsworth and Matthew Abraham and Elizabeth Dorfman and N. P. Armitage and Chunhan Feng and Antoine Georges and Olivier Gingras and Dominik Kiese and Steven A. Kivelson and Vadim Oganesyan and B. J. Ramshaw and Subir Sachdev and T. Senthil and J. M. Tranquada and Michael P. Brenner and Subhashini Venugopalan and Eun-Ah Kim},
journal= {arXiv preprint arXiv:2511.03782},
year = {2026}
}
备注
(v1) 9 pages, 4 figures, with 7-page supporting information. Accepted at the ICML 2025 workshop on Assessing World Models and the Explorations in AI Today workshop at ICML'25