English

Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl

Instrumentation and Methods for Astrophysics 2023-05-08 v3 Machine Learning Neural and Evolutionary Computing Symbolic Computation Data Analysis, Statistics and Probability

Abstract

PySR is an open-source library for practical symbolic regression, a type of machine learning which aims to discover human-interpretable symbolic models. PySR was developed to democratize and popularize symbolic regression for the sciences, and is built on a high-performance distributed back-end, a flexible search algorithm, and interfaces with several deep learning packages. PySR's internal search algorithm is a multi-population evolutionary algorithm, which consists of a unique evolve-simplify-optimize loop, designed for optimization of unknown scalar constants in newly-discovered empirical expressions. PySR's backend is the extremely optimized Julia library SymbolicRegression.jl, which can be used directly from Julia. It is capable of fusing user-defined operators into SIMD kernels at runtime, performing automatic differentiation, and distributing populations of expressions to thousands of cores across a cluster. In describing this software, we also introduce a new benchmark, "EmpiricalBench," to quantify the applicability of symbolic regression algorithms in science. This benchmark measures recovery of historical empirical equations from original and synthetic datasets.

Keywords

Cite

@article{arxiv.2305.01582,
  title  = {Interpretable Machine Learning for Science with PySR and SymbolicRegression.jl},
  author = {Miles Cranmer},
  journal= {arXiv preprint arXiv:2305.01582},
  year   = {2023}
}

Comments

24 pages, 5 figures, 3 tables. Feedback welcome. Paper source found at https://github.com/MilesCranmer/pysr_paper ; PySR at https://github.com/MilesCranmer/PySR ; SymbolicRegression.jl at https://github.com/MilesCranmer/SymbolicRegression.jl

R2 v1 2026-06-28T10:23:40.714Z