English
Related papers

Related papers: Aligning Subtitles in Sign Language Videos

200 papers

As the main means of communication for deaf people, sign language has a special grammatical order, so it is meaningful and valuable to develop a real-time translation system for sign language. In the research process, we added a TSM module…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Shengzhuo Wei , Yan Lan

This paper proposes an attentional network for the task of Continuous Sign Language Recognition. The proposed approach exploits co-independent streams of data to model the sign language modalities. These different channels of information…

Computer Vision and Pattern Recognition · Computer Science 2021-01-13 Fares Ben Slimane , Mohamed Bouguessa

Previous models for video captioning often use the output from a specific layer of a Convolutional Neural Network (CNN) as video features. However, the variable context-dependent semantics in the video may make it more appropriate to…

Computer Vision and Pattern Recognition · Computer Science 2017-11-20 Yunchen Pu , Martin Renqiang Min , Zhe Gan , Lawrence Carin

This paper presents a novel approach for temporal and semantic segmentation of edited videos into meaningful segments, from the point of view of the storytelling structure. The objective is to decompose a long video into more manageable…

Computer Vision and Pattern Recognition · Computer Science 2016-11-11 Lorenzo Baraldi , Costantino Grana , Rita Cucchiara

In this paper, we investigate the problem of unpaired video-to-video translation. Given a video in the source domain, we aim to learn the conditional distribution of the corresponding video in the target domain, without seeing any pairs of…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Kwanyong Park , Sanghyun Woo , Dahun Kim , Donghyeon Cho , In So Kweon

A persistent challenge in sign language video processing, including the task of sign to written language translation, is how we learn representations of sign language in an effective and efficient way that preserves the important attributes…

Computation and Language · Computer Science 2025-06-04 Shester Gueuwou , Xiaodan Du , Greg Shakhnarovich , Karen Livescu

We address the task of zero-shot video classification for extremely fine-grained actions (e.g., Windmill Dunk in basketball), where no video examples or temporal annotations are available for unseen classes. While image-language models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-21 Amir Aghdam , Vincent Tao Hu , Björn Ommer

We present a large-scale video subtitle translation dataset, BigVideo, to facilitate the study of multi-modality machine translation. Compared with the widely used How2 and VaTeX datasets, BigVideo is more than 10 times larger, consisting…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Liyan Kang , Luyang Huang , Ningxin Peng , Peihao Zhu , Zewei Sun , Shanbo Cheng , Mingxuan Wang , Degen Huang , Jinsong Su

Speech translation (ST) has lately received growing interest for the generation of subtitles without the need for an intermediate source language transcription and timing (i.e. captions). However, the joint generation of source captions and…

Computation and Language · Computer Science 2021-07-14 Alina Karakanta , Marco Gaido , Matteo Negri , Marco Turchi

The best summary of a long video differs among different people due to its highly subjective nature. Even for the same person, the best summary may change with time or mood. In this paper, we introduce the task of generating customized…

Computer Vision and Pattern Recognition · Computer Science 2018-03-05 Jinsoo Choi , Tae-Hyun Oh , In So Kweon

The goal of automatic Sign Language Production (SLP) is to translate spoken language to a continuous stream of sign language video at a level comparable to a human translator. If this was achievable, then it would revolutionise Deaf hearing…

Computer Vision and Pattern Recognition · Computer Science 2020-07-21 Ben Saunders , Necati Cihan Camgoz , Richard Bowden

Standard video and movie description tasks abstract away from person identities, thus failing to link identities across sentences. We propose a multi-sentence Identity-Aware Video Description task, which overcomes this limitation and…

Computer Vision and Pattern Recognition · Computer Science 2020-08-25 Jae Sung Park , Trevor Darrell , Anna Rohrbach

Video subtitles play a crucial role in short videos and movies, as they not only help models better understand video content but also support applications such as video translation and content retrieval. Existing video subtitle extraction…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Haiyang Yu , Mengyang Zhao , Jinghui Lu , Ke Niu , Yanjie Wang , Weijie Yin , Weitao Jia , Teng Fu , Yang Liu , Jun Liu , Hong Chen

Like spoken languages, a single sign language expression could correspond to multiple valid textual interpretations. Hence, learning a rigid one-to-one mapping for sign language translation (SLT) models might be inadequate, particularly in…

Multimedia · Computer Science 2024-10-28 Xin Shen , Lei Shen , Shaozu Yuan , Heming Du , Haiyang Sun , Xin Yu

We introduce a novel self-supervised pretext task for learning representations from audio-visual content. Prior work on audio-visual representation learning leverages correspondences at the video level. Approaches based on audio-visual…

Computer Vision and Pattern Recognition · Computer Science 2020-11-04 Pedro Morgado , Yi Li , Nuno Vasconcelos

Continuous sign language recognition (SLR) aims to translate a signing sequence into a sentence. It is very challenging as sign language is rich in vocabulary, while many among them contain similar gestures and motions. Moreover, it is…

Computer Vision and Pattern Recognition · Computer Science 2019-08-06 Zhaoyang Yang , Zhenmei Shi , Xiaoyong Shen , Yu-Wing Tai

We introduce a novel and inexpensive approach for the temporal alignment of speech to highly imperfect transcripts from automatic speech recognition (ASR). Transcripts are generated for extended lecture and presentation videos, which in…

Sound · Computer Science 2007-05-23 Alexander Haubold , John R. Kender

Millions of hearing impaired people around the world routinely use some variants of sign languages to communicate, thus the automatic translation of a sign language is meaningful and important. Currently, there are two sub-problems in Sign…

Computer Vision and Pattern Recognition · Computer Science 2025-09-16 Jie Huang , Wengang Zhou , Qilin Zhang , Houqiang Li , Weiping Li

A fitting soundtrack can help a video better convey its content and provide a better immersive experience. This paper introduces a novel approach utilizing self-supervised learning and contrastive learning to automatically recommend audio…

Multimedia · Computer Science 2025-03-10 Shimiao Liu , Alexander Lerch

In recent years, video semantic segmentation has made great progress with advanced deep neural networks. However, there still exist two main challenges \ie, information inconsistency and computation cost. To deal with the two difficulties,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-19 Jinming Su , Ruihong Yin , Shuaibin Zhang , Junfeng Luo
‹ Prev 1 4 5 6 7 8 10 Next ›