English

Fair and Calibrated Toxicity Detection with Robust Training and Abstention

Machine Learning 2026-05-15 v1

Abstract

Fairness in toxicity classification involves three integrated axes: ranking, calibration, and abstention. Training-time interventions and post-hoc safety mechanisms cannot be evaluated independently because the former determines the efficacy of the latter. We compare Empirical Risk Minimization (ERM), instance-level reweighting, and Group DRO across these axes, combined with temperature scaling, confidence-based abstention, and per-identity threshold optimization. Evaluation uses subgroup AUC, BPSN/BNSP AUC, error gaps, and per-subgroup Expected Calibration Error (ECE) with bootstrap CIs (n=1000n = 1000). We report four findings. (1) Calibration disparity is a hidden fairness violation. ERM has near-perfect aggregate calibration (0.0130.013) but is significantly miscalibrated across all identity subgroups (+0.029+0.029 to +0.134+0.134). (2) Training interventions reshape rather than eliminate disparity. Reweighted ERM improves ranking (BPSN AUC +0.06+0.06 to +0.12+0.12) but worsens the calibration-fairness gap by up to +0.232+0.232. Group DRO eliminates calibration disparity but only by becoming uniformly miscalibrated globally (ECE 0.1180.118). (3) Post-hoc methods inherit training failure modes. Temperature scaling fails because miscalibration is non-uniform. Confidence-based abstention works under ERM but breaks under DRO, where the risk-coverage curve rises with deferral. (4) Abstention itself is unfair. Confidence-based deferral helps background content far more than identity-mentioning content. We argue that SRAI fairness requires a multi-axis framework: methods that differ only in aggregate ranking can differ sharply in failure modes that determine real-world harm.

Keywords

Cite

@article{arxiv.2605.14074,
  title  = {Fair and Calibrated Toxicity Detection with Robust Training and Abstention},
  author = {Mokshit Surana},
  journal= {arXiv preprint arXiv:2605.14074},
  year   = {2026}
}
R2 v1 2026-07-22T07:11:06.792Z