Optimizing the extraction of information from redshift probability distribution functions
Abstract
Photometric redshifts are essential for large-scale structure analyses, yet extracting optimal point estimates and reliability measures from the probability distribution functions (PDZs) delivered by photo- pipelines remains an open challenge. We introduce turboPDZ, a machine-learning framework that optimizes both quantities directly from the PDZ. We apply the framework to PDZs from the three independent HSC-SSP PDR3 pipelines (DEmP, DNNz, Mizuki) across Wide and DUD layers. Each PDZ is compressed via PCA and combined with summary descriptors; a multilayer perceptron, optimized with Optuna under a composite objective, produces the optimized point estimate . A second network, trained in log-space and calibrated, yields the uncertainty , from which the reliability score is derived via percentile ranking. outperforms the catalog in and across all six pipeline-layer combinations. filters galaxies more efficiently than the catalog risk and confidence indicators, as measured by the area under the and versus retained-fraction curves. For Mizuki, the template-fitting pipeline, the catalog indicators fail dramatically, with AUC values up to ten times larger than those of , whereas correctly identifies unreliable objects across all redshift regimes. Feature-importance analysis reveals complementary patterns: point estimation is dominated by PCA components and location statistics, while reliability estimation depends on PCA components and peak statistics. The pipeline is survey-independent, publicly available at https://github.com/valerio-marra/turboPDZ, and trained models plus optimized quantities are released as a value-added catalog.
Cite
@article{arxiv.2607.26822,
title = {Optimizing the extraction of information from redshift probability distribution functions},
author = {Rodrigo Duarte and Valerio Marra},
journal= {arXiv preprint arXiv:2607.26822},
year = {2026}
}
Comments
14 pages, 11 figures