English

Hidden-State Privacy Has an Empty Middle

Machine Learning 2026-05-27 v2 Artificial Intelligence

Abstract

Of 1,5361{,}536 Gaussian release covariances we tested for single-layer hidden-state privacy, zero achieve both moderate utility and moderate privacy against an adaptive retrieval attacker. We prove a complementary Fisher-ball lower bound: every full-rank Gaussian release at O(1)O(1) Fisher utility admits a direction whose Mahalanobis signal grows linearly in hidden width, ruling out uniform Gaussian safety in the class and matching the empirical empty middle. The diagonal inverse-Fisher release Σdiag(K)=(2K/d)diag(1/Fii)\Sigma^\star_{\mathrm{diag}}(\mathcal{K}) = (2\mathcal{K}/d)\,\mathrm{diag}(1/F_{ii}) is the unique minimax-optimal diagonal mechanism at first-order KL budget K\mathcal{K} and the only release with worst-attacker top-1 0.001\le 0.001 at every point of a 32 model-layer grid, but it sits on a privacy/utility edge rather than filling the middle. A generalized-eigen mechanism reaching 13×13\times Pareto reduction under Euclidean retrieval collapses to 100%100\% top-1 under the adaptive Mahalanobis attacker, and a full-trajectory sequence inverter recovers 94%94\% of clean GPT-2 prefixes but 0%0\% under Σdiag\Sigma_{\mathrm{diag}}. A split-memory transformer trained from scratch reaches GMah[20,33]G_{\mathrm{Mah}} \in [20, 33] at 90M and maintains a 66--24×24\times advantage over same-budget GPT baselines from 30M to 1B at a fixed-token language-modeling loss penalty; pretrained models top out at 9.3. These results reframe hidden-state release from mechanism-design within the Gaussian class to architecture or release co-design.

Cite

@article{arxiv.2605.24042,
  title  = {Hidden-State Privacy Has an Empty Middle},
  author = {Alexander Okezue Bell},
  journal= {arXiv preprint arXiv:2605.24042},
  year   = {2026}
}

Comments

74 pages, 61 figures

R2 v1 2026-07-22T07:29:06.189Z