How Open Must Language Models be to Enable Reliable Scientific Inference?
Computation and Language
2026-05-21 v2 Artificial Intelligence
Abstract
How does the extent to which a model is open or closed impact the scientific inferences that can be drawn from research that involves it? In this paper, we analyze how restrictions on information about model construction and deployment threaten reliable inference. We argue that current closed models are generally ill-suited for scientific purposes, with some notable exceptions, and discuss ways in which the issues they present to reliable inference can be resolved or mitigated. We recommend that when models are used in research, potential threats to inference should be systematically identified along with the steps taken to mitigate them, and that specific justifications for model selection should be provided.
Cite
@article{arxiv.2603.26539,
title = {How Open Must Language Models be to Enable Reliable Scientific Inference?},
author = {James A. Michaelov and Catherine Arnett and Tyler A. Chang and Pamela D. Rivière and Samuel M. Taylor and Cameron R. Jones and Sean Trott and Roger P. Levy and Benjamin K. Bergen and Micah Altman},
journal= {arXiv preprint arXiv:2603.26539},
year = {2026}
}