English

SLVideo: A Sign Language Video Moment Retrieval Framework

Computer Vision and Pattern Recognition 2024-11-07 v2 Artificial Intelligence

Abstract

SLVideo is a video moment retrieval system for Sign Language videos that incorporates facial expressions, addressing this gap in existing technology. The system extracts embedding representations for the hand and face signs from video frames to capture the signs in their entirety, enabling users to search for a specific sign language video segment with text queries. A collection of eight hours of annotated Portuguese Sign Language videos is used as the dataset, and a CLIP model is used to generate the embeddings. The initial results are promising in a zero-shot setting. In addition, SLVideo incorporates a thesaurus that enables users to search for similar signs to those retrieved, using the video segment embeddings, and also supports the edition and creation of video sign language annotations. Project web page: https://novasearch.github.io/SLVideo/

Keywords

Cite

@article{arxiv.2407.15668,
  title  = {SLVideo: A Sign Language Video Moment Retrieval Framework},
  author = {Gonçalo Vinagre Martins and João Magalhães and Afonso Quinaz and Carla Viegas and Sofia Cavaco},
  journal= {arXiv preprint arXiv:2407.15668},
  year   = {2024}
}

Comments

4 pages, 1 figure, 1 table

R2 v1 2026-06-28T17:49:33.919Z