English

Beyond the Fold: Quantifying Split-Level Noise and the Case for Leave-One-Dataset-Out AU Evaluation

Computer Vision and Pattern Recognition 2026-04-03 v1

Abstract

Subject-exclusive cross-validation is the standard evaluation protocol for facial Action Unit (AU) detection, yet reported improvements are often small. We show that cross-validation itself introduces measurable stochastic variance. On BP4D+, repeated 3-fold subject-exclusive splits produce an empirical noise floor of ±0.065\pm 0.065 in average F1, with substantially larger variation for low-prevalence AUs. Operating-point metrics such as F1 fluctuate more than threshold-independent measures such as AUC, and model ranking can change under different fold assignments. We further evaluate cross-dataset robustness using a Leave-One-Dataset-Out (LODO) protocol across five AU datasets. LODO removes partition randomness and exposes domain-level instability that is not visible under single-dataset cross-validation. Together, these results suggest that gains often reported in cross-fold validation may fall within protocol variance. Leave-one-dataset-out cross-validation yields more stable and interpretable findings

Cite

@article{arxiv.2604.02162,
  title  = {Beyond the Fold: Quantifying Split-Level Noise and the Case for Leave-One-Dataset-Out AU Evaluation},
  author = {Saurabh Hinduja and Gurmeet Kaur and Maneesh Bilalpur and Jeffrey Cohn and Shaun Canavan},
  journal= {arXiv preprint arXiv:2604.02162},
  year   = {2026}
}

Comments

CVPR 2026

R2 v1 2026-07-01T11:51:14.606Z