中文
相关论文

相关论文: SLVideo: A Sign Language Video Moment Retrieval Fr…

200 篇论文

Textual overlays are often used in social media videos as people who watch them without the sound would otherwise miss essential information conveyed in the audio stream. This is why extraction of those overlays can serve as an important…

计算机视觉与模式识别 · 计算机科学 2018-05-02 Adam Słucki , Tomasz Trzcinski , Adam Bielski , Paweł Cyrta

For video-text retrieval, the use of CLIP has been a de facto choice. Since CLIP provides only image and text encoders, this consensus has led to a biased paradigm that entirely ignores the sound track of videos. While several attempts have…

计算机视觉与模式识别 · 计算机科学 2026-03-12 Ruixiang Zhao , Zhihao Xu , Bangxiang Lan , Zijie Xin , Jingyu Liu , Xirong Li

We propose a real time deep learning framework for video-based facial expression capture. Our process uses a high-end facial capture pipeline based on FACEGOOD to capture facial expression. We train a convolutional neural network to produce…

计算机视觉与模式识别 · 计算机科学 2021-11-16 Hongwei Xu , Leijia Dai , Jianxing Fu , Xiangyuan Wang , Quanwei Wang

Retrieving adverbs that describe an action in a video poses a crucial step towards fine-grained video understanding. We propose a framework for video-to-adverb retrieval (and vice versa) that aligns video embeddings with their matching…

计算机视觉与模式识别 · 计算机科学 2023-09-27 Thomas Hummel , Otniel-Bogdan Mercea , A. Sophia Koepke , Zeynep Akata

The American Sign Language Linguistic Research Project (ASLLRP) provides Internet access to high-quality ASL video data, generally including front and side views and a close-up of the face. The manual and non-manual components of the…

计算与语言 · 计算机科学 2022-01-21 Carol Neidle , Augustine Opoku , Dimitris Metaxas

Disciplines such as business process management and process mining aid organizations by discovering insights about processes on the basis of recorded event data. However, an obstacle to process analysis is data multi-modality: for instance,…

计算机视觉与模式识别 · 计算机科学 2026-04-27 Marco Pegoraro , Jonas Seng , Dustin Heller , Wil M. P. van der Aalst , Kristian Kersting

In this paper, we present a framework that jointly retrieves and spatiotemporally highlights actions in videos by enhancing current deep cross-modal retrieval methods. Our work takes on the novel task of action highlighting, which…

计算机视觉与模式识别 · 计算机科学 2020-05-20 Seito Kasai , Yuchi Ishikawa , Masaki Hayashi , Yoshimitsu Aoki , Kensho Hara , Hirokatsu Kataoka

The recent success of the CLIP model has shown its potential to be applied to a wide range of vision and language tasks. However this only establishes embedding space relationship of language to images, not to the video domain. In this…

计算机视觉与模式识别 · 计算机科学 2023-04-11 Phani Krishna Uppala , Abhishek Bamotra , Shriti Priya , Vaidehi Joshi

There has been a growing interest in the task of generating sound for silent videos, primarily because of its practicality in streamlining video post-production. However, existing methods for video-sound generation attempt to directly…

多媒体 · 计算机科学 2024-04-04 Zhifeng Xie , Shengye Yu , Qile He , Mengtian Li

We present SignCLIP, which re-purposes CLIP (Contrastive Language-Image Pretraining) to project spoken language text and sign language videos, two classes of natural languages of distinct modalities, into the same space. SignCLIP is an…

计算与语言 · 计算机科学 2024-10-08 Zifan Jiang , Gerard Sant , Amit Moryossef , Mathias Müller , Rico Sennrich , Sarah Ebling

We propose a high-level concept word detector that can be integrated with any video-to-language models. It takes a video as input and generates a list of concept words as useful semantic priors for language generation models. The proposed…

计算机视觉与模式识别 · 计算机科学 2017-07-26 Youngjae Yu , Hyungjin Ko , Jongwook Choi , Gunhee Kim

As a natural extension of the image synthesis task, video synthesis has attracted a lot of interest recently. Many image synthesis works utilize class labels or text as guidance. However, neither labels nor text can provide explicit…

计算机视觉与模式识别 · 计算机科学 2022-11-18 Yuren Cong , Jinhui Yi , Bodo Rosenhahn , Michael Ying Yang

Videos are a rich source for self-supervised learning (SSL) of visual representations due to the presence of natural temporal transformations of objects. However, current methods typically randomly sample video clips for learning, which…

计算机视觉与模式识别 · 计算机科学 2022-09-30 Brian Chen , Ramprasaath R. Selvaraju , Shih-Fu Chang , Juan Carlos Niebles , Nikhil Naik

Sign language recognition is a challenging and often underestimated problem comprising multi-modal articulators (handshape, orientation, movement, upper body and face) that integrate asynchronously on multiple streams. Learning powerful…

计算机视觉与模式识别 · 计算机科学 2019-11-22 Hamid Reza Vaezi Joze , Oscar Koller

Surgical phase recognition from video is a technology that automatically classifies the progress of a surgical procedure and has a wide range of potential applications, including real-time surgical support, optimization of medical…

计算机视觉与模式识别 · 计算机科学 2025-05-21 Satoshi Kondo

Sign language is the window for people differently-abled to express their feelings as well as emotions. However, it remains challenging for people to learn sign language in a short time. To address this real-world challenge, in this work,…

计算机视觉与模式识别 · 计算机科学 2022-07-11 Yucheng Suo , Zhedong Zheng , Xiaohan Wang , Bang Zhang , Yi Yang

Recent studies have shown promising results on utilizing large pre-trained image-language models for video question answering. While these image-language models can efficiently bootstrap the representation learning of video-language models,…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Shoubin Yu , Jaemin Cho , Prateek Yadav , Mohit Bansal

This study presents TSLFormer, a light and robust word-level Turkish Sign Language (TSL) recognition model that treats sign gestures as ordered, string-like language. Instead of using raw RGB or depth videos, our method only works with 3D…

计算与语言 · 计算机科学 2025-06-19 Kutay Ertürk , Furkan Altınışık , İrem Sarıaltın , Ömer Nezih Gerek

Sign Language Translation (SLT) is a challenging task that aims to translate sign videos into spoken language. Inspired by the strong translation capabilities of large language models (LLMs) that are trained on extensive multilingual text…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Jia Gong , Lin Geng Foo , Yixuan He , Hossein Rahmani , Jun Liu

Learning multimodal video understanding typically relies on datasets comprising video clips paired with manually annotated captions. However, this becomes even more challenging when dealing with long-form videos, lasting from minutes to…

计算机视觉与模式识别 · 计算机科学 2025-03-12 Soumya Shamarao Jahagirdar , Jayasree Saha , C V Jawahar