超越隐藏层操控:语义感知的逻辑干预去偏方法
机器学习
2025-10-29 v1 人工智能
摘要
我们提出了 Static 和 Dynamic 两种零样本逻辑层去偏方法。Dynamic 方法通过最小的流畅度损失将偏见降低最高可达 70%。逻辑干预优于隐藏层方法。我们表明,语义感知的逻辑干预对于去偏对齐型大语言模型是稳定且有效的。
引用
@article{arxiv.2510.23650,
title = {Beyond Hidden-Layer Manipulation: Semantically-Aware Logit Interventions for Debiasing LLMs},
author = {Wei Xia},
journal= {arXiv preprint arXiv:2510.23650},
year = {2025}
}