Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference
Abstract
Hong-Ou-Mandel (HOM) interference-based optical neural networks can offer complexity advantages on benchmark learning tasks, but conventional readout compresses the coincidence spectrum into a single scalar, limiting its use in complex settings such as continuous-action reinforcement learning. Here we introduce a spectrum-resolved HOM (SR-HOM) architecture that promotes the photons' spectral degrees of freedom to a trainable computational resource and use it to construct a compact optical actor-critic agent. Diagonal spectral responses generate continuous actions, while higher-order spectral correlations provide nonlinear state-action features for value estimation. Across five continuous-control benchmarks, SR-HOM outperforms parameter-matched multilayer-perceptron baselines, including a improvement in sample efficiency and a increase in best 100-episode moving-average return for LunarLanderContinuous-v3. Applied to online calibration of drifted tunable-coupler CZ and iSWAP gates for transmon qubits, simulations show it restores fidelities to and respectively, exceeding of their drift-free calibrated values.
Cite
@article{arxiv.2607.26438,
title = {Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference},
author = {Shaojun Wu and Jiahua Xu and Shan Jin and Zhen Yang and Yifang Xu and Chenglong You and Guangwei Deng and Luyan Sun and Chang-Ling Zou and Xiaoting Wang},
journal= {arXiv preprint arXiv:2607.26438},
year = {2026}
}