English

Scaling up sign spotting through sign language dictionaries

Computer Vision and Pattern Recognition 2022-05-10 v1

Abstract

The focus of this work is sign spotting\textit{sign spotting} - given a video of an isolated sign, our task is to identify whether\textit{whether} and where\textit{where} it has been signed in a continuous, co-articulated sign language video. To achieve this sign spotting task, we train a model using multiple types of available supervision by: (1) watching\textit{watching} existing footage which is sparsely labelled using mouthing cues; (2) reading\textit{reading} associated subtitles (readily available translations of the signed content) which provide additional weak-supervision\textit{weak-supervision}; (3) looking up\textit{looking up} words (for which no co-articulated labelled examples are available) in visual sign language dictionaries to enable novel sign spotting. These three tasks are integrated into a unified learning framework using the principles of Noise Contrastive Estimation and Multiple Instance Learning. We validate the effectiveness of our approach on low-shot sign spotting benchmarks. In addition, we contribute a machine-readable British Sign Language (BSL) dictionary dataset of isolated signs, BSLDict, to facilitate study of this task. The dataset, models and code are available at our project page.

Keywords

Cite

@article{arxiv.2205.04152,
  title  = {Scaling up sign spotting through sign language dictionaries},
  author = {Gül Varol and Liliane Momeni and Samuel Albanie and Triantafyllos Afouras and Andrew Zisserman},
  journal= {arXiv preprint arXiv:2205.04152},
  year   = {2022}
}

Comments

Appears in: 2022 International Journal of Computer Vision (IJCV). 25 pages. arXiv admin note: substantial text overlap with arXiv:2010.04002