English

Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs

Machine Learning 2025-10-29 v1 Artificial Intelligence

Abstract

We proposed Static and Dynamic -- two zero-shot logits-layer debiasing methods. Dynamic reduces bias by up to 70% with minimal fluency loss. Logits intervention outperforms hidden-layer approaches. We show semantic-aware logits intervention is stable and effective for debiasing aligned LLMs.

Cite

@article{arxiv.2510.23650,
  title  = {Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs},
  author = {Wei Xia},
  journal= {arXiv preprint arXiv:2510.23650},
  year   = {2025}
}
R2 v1 2026-07-01T07:08:12.630Z