English

SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines

Computation and Language 2025-12-16 v3

Abstract

We present a scientific reasoning foundation model that aligns natural language with heterogeneous scientific representations. The model is pretrained on a 206B-token corpus spanning scientific text, pure sequences, and sequence-text pairs, then aligned via SFT on 40M instructions, annealed cold-start bootstrapping to elicit long-form chain-of-thought, and reinforcement learning with task-specific reward shaping, which instills deliberate scientific reasoning. It supports four capability families, covering up to 103 tasks across workflows: (i) faithful translation between text and scientific formats, (ii) text/knowledge extraction, (iii) property prediction, (iv) property classification, (v) unconditional and conditional sequence generation and design. Compared with specialist systems, our approach broadens instruction coverage, improves cross-domain generalization, and enhances fidelity. We detail data curation and training and show that cross-discipline learning strengthens transfer and downstream reliability. The model, instruct tuning datasets and the evaluation code are open-sourced at https://huggingface.co/SciReason and https://github.com/open-sciencelab/SciReason.

Keywords

Cite

@article{arxiv.2509.21320,
  title  = {SciReasoner: Laying the Scientific Reasoning Ground Across Disciplines},
  author = {Yizhou Wang and Chen Tang and Han Deng and Jiabei Xiao and Jiaqi Liu and Jianyu Wu and Jun Yao and Pengze Li and Encheng Su and Lintao Wang and Guohang Zhuang and Yuchen Ren and Ben Fei and Ming Hu and Xin Chen and Dongzhan Zhou and Junjun He and Xiangyu Yue and Zhenfei Yin and Jiamin Wu and Qihao Zheng and Yuhao Zhou and Huihui Xu and Chenglong Ma and Yan Lu and Wenlong Zhang and Chunfeng Song and Philip Torr and Shixiang Tang and Xinzhu Ma and Wanli Ouyang and Lei Bai},
  journal= {arXiv preprint arXiv:2509.21320},
  year   = {2025}
}

Comments

technical report