We explore the ability of GPT-4 to perform ad-hoc schema based information extraction from scientific literature. We assess specifically whether it can, with a basic prompting approach, replicate two existing material science datasets, given the manuscripts from which they were originally manually extracted. We employ materials scientists to perform a detailed manual error analysis to assess where the model struggles to faithfully extract the desired information, and draw on their insights to suggest research directions to address this broadly important task.
@article{arxiv.2406.05348,
title = {Toward Reliable Ad-hoc Scientific Information Extraction: A Case Study on Two Materials Datasets},
author = {Satanu Ghosh and Neal R. Brodnik and Carolina Frey and Collin Holgate and Tresa M. Pollock and Samantha Daly and Samuel Carton},
journal= {arXiv preprint arXiv:2406.05348},
year = {2025}
}
Comments
LLM for information extraction. Update on 12/11/2024: We added some relevant literature that we missed in the previous version of the paper. Update on 05/25/2025: We changed the metadata