English

Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference

Quantum Physics 2026-07-29 v1

Abstract

Hong-Ou-Mandel (HOM) interference-based optical neural networks can offer complexity advantages on benchmark learning tasks, but conventional readout compresses the coincidence spectrum into a single scalar, limiting its use in complex settings such as continuous-action reinforcement learning. Here we introduce a spectrum-resolved HOM (SR-HOM) architecture that promotes the photons' spectral degrees of freedom to a trainable computational resource and use it to construct a compact optical actor-critic agent. Diagonal spectral responses generate continuous actions, while higher-order spectral correlations provide nonlinear state-action features for value estimation. Across five continuous-control benchmarks, SR-HOM outperforms parameter-matched multilayer-perceptron baselines, including a 4.4×4.4\times improvement in sample efficiency and a 74.0%74.0\% increase in best 100-episode moving-average return for LunarLanderContinuous-v3. Applied to online calibration of drifted tunable-coupler CZ and iSWAP gates for transmon qubits, simulations show it restores fidelities to 0.99170.9917 and 0.99520.9952 respectively, exceeding 99.8%99.8\% of their drift-free calibrated values.

Cite

@article{arxiv.2607.26438,
  title  = {Quantum Optical Reinforcement Learning via Spectrum-Resolved Hong-Ou-Mandel Interference},
  author = {Shaojun Wu and Jiahua Xu and Shan Jin and Zhen Yang and Yifang Xu and Chenglong You and Guangwei Deng and Luyan Sun and Chang-Ling Zou and Xiaoting Wang},
  journal= {arXiv preprint arXiv:2607.26438},
  year   = {2026}
}