English

Optimal Convergence Rates of Deep Neural Network Classifiers

Machine Learning 2025-11-24 v2 Machine Learning

Abstract

In this paper, we study the binary classification problem on [0,1]d[0,1]^d under the Tsybakov noise condition (with exponent s[0,]s \in [0,\infty]) and the compositional assumption. This assumption requires the conditional class probability function of the data distribution to be the composition of q+1q+1 vector-valued multivariate functions, where each component function is either a maximum value function or a H\"{o}lder-β\beta smooth function that depends only on dd_* of its input variables. Notably, dd_* can be significantly smaller than the input dimension dd. We prove that, under these conditions, the optimal convergence rate for the excess 0-1 risk of classifiers is (1n)β(1β)qds+1+(1+1s+1)β(1β)q\left( \frac{1}{n} \right)^{\frac{\beta\cdot(1\wedge\beta)^q}{{\frac{d_*}{s+1}+(1+\frac{1}{s+1})\cdot\beta\cdot(1\wedge\beta)^q}}}, which is independent of the input dimension dd. Additionally, we demonstrate that ReLU deep neural networks (DNNs) trained with hinge loss can achieve this optimal convergence rate up to a logarithmic factor. This result provides theoretical justification for the excellent performance of ReLU DNNs in practical classification tasks, particularly in high-dimensional settings. The generalized approach is of independent interest.

Keywords

Cite

@article{arxiv.2506.14899,
  title  = {Optimal Convergence Rates of Deep Neural Network Classifiers},
  author = {Zihan Zhang and Lei Shi and Ding-Xuan Zhou},
  journal= {arXiv preprint arXiv:2506.14899},
  year   = {2025}
}
R2 v1 2026-07-01T03:22:38.118Z