English

Decorrelate Irrelevant, Purify Relevant: Overcome Textual Spurious Correlations from a Feature Perspective

Computation and Language 2022-09-14 v2 Artificial Intelligence Information Theory math.IT

Abstract

Natural language understanding (NLU) models tend to rely on spurious correlations (i.e., dataset bias) to achieve high performance on in-distribution datasets but poor performance on out-of-distribution ones. Most of the existing debiasing methods often identify and weaken these samples with biased features (i.e., superficial surface features that cause such spurious correlations). However, down-weighting these samples obstructs the model in learning from the non-biased parts of these samples. To tackle this challenge, in this paper, we propose to eliminate spurious correlations in a fine-grained manner from a feature space perspective. Specifically, we introduce Random Fourier Features and weighted re-sampling to decorrelate the dependencies between features to mitigate spurious correlations. After obtaining decorrelated features, we further design a mutual-information-based method to purify them, which forces the model to learn features that are more relevant to tasks. Extensive experiments on two well-studied NLU tasks demonstrate that our method is superior to other comparative approaches.

Keywords

Cite

@article{arxiv.2202.08048,
  title  = {Decorrelate Irrelevant, Purify Relevant: Overcome Textual Spurious Correlations from a Feature Perspective},
  author = {Shihan Dou and Rui Zheng and Ting Wu and SongYang Gao and Junjie Shan and Qi Zhang and Yueming Wu and Xuanjing Huang},
  journal= {arXiv preprint arXiv:2202.08048},
  year   = {2022}
}

Comments

Accepted as a long paper at COLING 2022 (Oral)

R2 v1 2026-06-24T09:40:53.984Z