English

Bilevel Learning with Inexact Stochastic Gradients

Optimization and Control 2025-05-20 v2 Machine Learning

Abstract

Bilevel learning has gained prominence in machine learning, inverse problems, and imaging applications, including hyperparameter optimization, learning data-adaptive regularizers, and optimizing forward operators. The large-scale nature of these problems has led to the development of inexact and computationally efficient methods. Existing adaptive methods predominantly rely on deterministic formulations, while stochastic approaches often adopt a doubly-stochastic framework with impractical variance assumptions, enforces a fixed number of lower-level iterations, and requires extensive tuning. In this work, we focus on bilevel learning with strongly convex lower-level problems and a nonconvex sum-of-functions in the upper-level. Stochasticity arises from data sampling in the upper-level which leads to inexact stochastic hypergradients. We establish their connection to state-of-the-art stochastic optimization theory for nonconvex objectives. Furthermore, we prove the convergence of inexact stochastic bilevel optimization under mild assumptions. Our empirical results highlight significant speed-ups and improved generalization in imaging tasks such as image denoising and deblurring in comparison with adaptive deterministic bilevel methods.

Keywords

Cite

@article{arxiv.2412.12049,
  title  = {Bilevel Learning with Inexact Stochastic Gradients},
  author = {Mohammad Sadegh Salehi and Subhadip Mukherjee and Lindon Roberts and Matthias J. Ehrhardt},
  journal= {arXiv preprint arXiv:2412.12049},
  year   = {2025}
}

Comments

Accepted to the 10th International Conference on Scale Space and Variational Methods in Computer Vision (SSVM 2025)

R2 v1 2026-06-28T20:37:29.560Z