English

Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval

Computer Vision and Pattern Recognition 2026-07-10 v1

Abstract

Sign Language Retrieval (SLRet) enables efficient access to sign language content but remains fragile in fine-grained scenarios where visually similar signs must be distinguished. We show that this limitation does not stem from model capacity, but from ineffective hard negative supervision. Specifically, we formulate fine-grained retrieval failures as a negative distribution mismatch: semantically distinct yet visually confusable signs are rarely treated as hard negatives, while existing text-based mining strategies fail to capture such visual ambiguity. To address this issue, we propose Sign-Aware Hard Negative Mining (SAN), which constructs hard negatives based on visual confusability in the sign embedding space rather than linguistic similarity. Experiments on PHOENIX-2014T demonstrate that SAN substantially improves fine-grained retrieval performance while preserving coarse-grained accuracy, highlighting the importance of aligning negative supervision with visual ambiguity in sign language retrieval.

Keywords

Cite

@article{arxiv.2607.09263,
  title  = {Semantic Hardness Is Not Visual Hardness: Sign-Aware Hard Negative Mining for Sign Language Retrieval},
  author = {Junmyeong Lee and Chan Hur and ChangSu Choi and Sukmin Cho and Fitsum Gaim and Eui Jun Hwang and Hoyun Song and KyungTae Lim},
  journal= {arXiv preprint arXiv:2607.09263},
  year   = {2026}
}

Comments

Accepted to ACL 2026 main