English

PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning

Computation and Language 2026-07-08 v1 Machine Learning

Abstract

One-shot pruning methods like Wanda and SparseGPT apply the same sparsity ratio to every layer of a transformer, ignoring known variation in layer importance. We propose PALS (Percentile-Aware Layerwise Sparsity), which adjusts per-layer sparsity based on the 99th percentile of activation magnitudes, bounded to ±5%\pm 5\% around the target ratio. On LLaMA-2-7B at 50\% sparsity, PALS achieves 10.96 WikiText-2 perplexity versus 12.92 for uniform Wanda (mean over 9 runs, p<0.001p < 0.001). The benefit is architecture-dependent: LLaMA-3-8B shows marginal gains and Mistral-7B shows none. We also find that gradient-based allocation -- the seemingly more principled approach -- produces results worse than random, suggesting that gradient magnitude does not predict the impact of discrete weight removal. PALS adds negligible cost to the pruning pipeline and requires no fine-tuning.

Cite

@article{arxiv.2607.07557,
  title  = {PALS: Percentile-Aware Layerwise Sparsity for LLM Pruning},
  author = {Yazdan Jamshidi and Alexey Shvets},
  journal= {arXiv preprint arXiv:2607.07557},
  year   = {2026}
}