A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models
Abstract
Activation steering via sparse autoencoders (SAEs) enables behavioral control of large language models without task-specific fine-tuning, but standard methods apply the steering signal at every generated token, incurring constant per-token perturbation that risks degrading fluency. We ask: is dense intervention necessary? We introduce Stochastic Token Steering (STS), which gates each token independently with probability , and Stochastic Block Steering (SBS), which gates a leading window once per sequence; neither requires a reward model or learned gating policy. Across two model families and two behavioral tasks, steering only 50% of the tokens recovers most of the dense-steering effect while preserving fluency, and steering as few as 30% surpasses prompt-based control. The optimal steering magnitude scales inversely with the intervention ratio, revealing that SAE-mediated control is rate-limited: the behavioral outcome depends on cumulative signal dosage across a sequence.
Keywords
Cite
@article{arxiv.2607.05615,
title = {A Coin Flip Per Token: Bernoulli Sparse Steering of Large Language Models},
author = {Nima Eshraghi and Lovedeep Gondara and Yuqing Huang and Sagarika Suresh and Leizer Teran and Jithin Pradeep and Xiaotong Xu and Fanny Chevalier},
journal= {arXiv preprint arXiv:2607.05615},
year = {2026}
}