English

Revocable Learned State via Process Sidecars

Machine Learning 2026-06-29 v1 Computation and Language Cryptography and Security

Abstract

Language models are often adapted in stages: a public skill phase, a private memory phase, and a later safety phase that learns to refuse outputs tied to the remembered entities. Revoking the memory after the safety phase is not the same problem as subtracting the memory update: the later safety optimizer has transported the memory direction. We introduce process sidecars, a two-coefficient edit family θ^(λ,γ)=θAMSλΔMγR^SM\hat{\theta}(\lambda,\gamma)=\theta_{\mathrm{AMS}}-\lambda\Delta_{\mathrm{M}}-\gamma\hat{R}_{\mathrm{S}\leftarrow\mathrm{M}}, with R^SM=J^S,ε(ΔM)ΔM\hat{R}_{\mathrm{S}\leftarrow\mathrm{M}}=\hat{J}_{\mathrm{S},\varepsilon}(\Delta_{\mathrm{M}})-\Delta_{\mathrm{M}}, where J^S,ε\hat{J}_{\mathrm{S},\varepsilon} is a centered secant through the realized future AdamW safety-training process. The implementation uses ε=1\varepsilon=1 at the natural memory-edit scale; it reuses θAMS\theta_{\mathrm{AMS}} as the positive endpoint and computes one additional safety trace at θAΔM\theta_{\mathrm{A}}-\Delta_{\mathrm{M}}. We prove two things. First, the exact sidecar, using the true transported direction RSMR_{\mathrm{S}\leftarrow\mathrm{M}} rather than the secant estimate, at (λ,γ)=(1,1)(\lambda,\gamma)=(1,1) recovers the counterfactual safety-only oracle θAS\theta_{\mathrm{AS}} up to second order; the proof treats AdamW as an augmented-state map over parameters, first moments, and second moments. Second, this process information is necessary: whenever future safety training bends the memory direction, every scalar task-arithmetic edit leaves first-order counterfactual error, while the process-sidecar edit is second-order accurate. Across three models, the validation-selected 2D edit improves held-out refusal closure over naive task arithmetic in all trials, and over the γ=λ\gamma=\lambda process-JVP subfamily, the diagonal slice of the cached 2D grid, in all paired trials.

Cite

@article{arxiv.2606.30788,
  title  = {Revocable Learned State via Process Sidecars},
  author = {John Sweeney},
  journal= {arXiv preprint arXiv:2606.30788},
  year   = {2026}
}

Comments

23 pages, 2 figures, 6 tables