English

Scaling Scientific Discovery Environments for Turn-Level Agentic RL

Artificial Intelligence 2026-07-31 v1

Abstract

Large language model agents have shown promising capabilities in data-driven scientific discovery tasks, where an agent interacts with an execution environment and produces a statistical claim. Long-horizon scientific analysis remains constrained by the lack of process supervised environments over real-world scientific data. This paper introduces SciDisco, a scalable framework for training Scientific Discovery agents in process-verifiable environments. SciTh\`eque compiles hypotheses, datasets, hidden evidence graphs, and verifiers into task environments where analytical progress can be checked during interaction. DAG-grounded trajectory synthesis uses these environments to construct verifier-filtered multi-turn demonstrations. DiscoPO then uses the environment as the source of training signal, assigning turn-level credit to actions that produce verifiable analytical evidence. Experiments show that SciDisco-14B reaches state-of-the-art on hypothesis-driven scientific data analysis benchmarks.

Cite

@article{arxiv.2607.28990,
  title  = {Scaling Scientific Discovery Environments for Turn-Level Agentic RL},
  author = {Yucheng Xu and Keyi Zhang and Yuyang Yu and Min Zhang and Shiyuan Meng and Pei Chu and Zhongying Tu},
  journal= {arXiv preprint arXiv:2607.28990},
  year   = {2026}
}