English

The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing

Computation and Language 2024-01-17 v1

Abstract

The output tendencies of Pre-trained Language Models (PLM) vary markedly before and after Fine-Tuning (FT) due to the updates to the model parameters. These divergences in output tendencies result in a gap in the social biases of PLMs. For example, there exits a low correlation between intrinsic bias scores of a PLM and its extrinsic bias scores under FT-based debiasing methods. Additionally, applying FT-based debiasing methods to a PLM leads to a decline in performance in downstream tasks. On the other hand, PLMs trained on large datasets can learn without parameter updates via In-Context Learning (ICL) using prompts. ICL induces smaller changes to PLMs compared to FT-based debiasing methods. Therefore, we hypothesize that the gap observed in pre-trained and FT models does not hold true for debiasing methods that use ICL. In this study, we demonstrate that ICL-based debiasing methods show a higher correlation between intrinsic and extrinsic bias scores compared to FT-based methods. Moreover, the performance degradation due to debiasing is also lower in the ICL case compared to that in the FT case.

Keywords

Cite

@article{arxiv.2401.08511,
  title  = {The Gaps between Pre-train and Downstream Settings in Bias Evaluation and Debiasing},
  author = {Masahiro Kaneko and Danushka Bollegala and Timothy Baldwin},
  journal= {arXiv preprint arXiv:2401.08511},
  year   = {2024}
}
R2 v1 2026-06-28T14:18:14.951Z