English
Related papers

Related papers: LRS3-TED: a large-scale dataset for visual speech …

200 papers

Closed-Set speaker identification aims to assign a speech utterance to one of a predefined set of enrolled speakers and requires robust modeling of speaker-specific characteristics across multiple temporal scales. While recent deep learning…

Sound · Computer Science 2026-05-11 Yassin Terraf , Youssef Iraqi

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent…

Sound · Computer Science 2020-11-05 Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Distant supervision (DS) is a well established technique for creating large-scale datasets for relation extraction (RE) without using human annotations. However, research in DS-RE has been mostly limited to the English language.…

Computation and Language · Computer Science 2021-04-20 Abhyuday Bhartiya , Kartikeya Badola , Mausam

Attributes of sound inherent to objects can provide valuable cues to learn rich representations for object detection and tracking. Furthermore, the co-occurrence of audiovisual events in videos can be exploited to localize objects over the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-05 Francisco Rivera Valverde , Juana Valeria Hurtado , Abhinav Valada

Visual speech recognition (VSR), commonly known as lip reading, has garnered significant attention due to its wide-ranging practical applications. The advent of deep learning techniques and advancements in hardware capabilities have…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Bowen Hao , Dongliang Zhou , Xiaojie Li , Xingyu Zhang , Liang Xie , Jianlong Wu , Erwei Yin

As the volume of long-form spoken-word content such as podcasts explodes, many platforms desire to present short, meaningful, and logically coherent segments extracted from the full content. Such segments can be consumed by users to sample…

Computation and Language · Computer Science 2021-12-13 Elise Jing , Kristiana Schneck , Dennis Egan , Scott A. Waterman

3D facial animation has attracted considerable attention due to its extensive applications in the multimedia field. Audio-driven 3D facial animation has been widely explored with promising results. However, multi-modal 3D facial animation,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-11 Sijing Wu , Yunhao Li , Yichao Yan , Huiyu Duan , Ziwei Liu , Guangtao Zhai

Most existing datasets for sound event recognition (SER) are relatively small and/or domain-specific, with the exception of AudioSet, based on over 2M tracks from YouTube videos and encompassing over 500 sound classes. However, AudioSet is…

Sound · Computer Science 2022-04-26 Eduardo Fonseca , Xavier Favory , Jordi Pons , Frederic Font , Xavier Serra

Even for better-studied sign languages like American Sign Language (ASL), data is the bottleneck for machine learning research. The situation is worse yet for the many other sign languages used by Deaf/Hard of Hearing communities around the…

Computation and Language · Computer Science 2024-07-17 Garrett Tanzer , Biao Zhang

An increasing number of datasets contain multiple views, such as video, sound and automatic captions. A basic challenge in representation learning is how to leverage multiple views to learn better representations. This is further…

Machine Learning · Computer Science 2019-03-04 Nils Holzenberger , Shruti Palaskar , Pranava Madhyastha , Florian Metze , Raman Arora

We introduce Replay, a collection of multi-view, multi-modal videos of humans interacting socially. Each scene is filmed in high production quality, from different viewpoints with several static cameras, as well as wearable action cameras,…

Computer Vision and Pattern Recognition · Computer Science 2023-07-25 Roman Shapovalov , Yanir Kleiman , Ignacio Rocco , David Novotny , Andrea Vedaldi , Changan Chen , Filippos Kokkinos , Ben Graham , Natalia Neverova

Overlapped Speech Detection (OSD) is an important part of speech applications involving analysis of multi-party conversations. However, most of existing OSD systems are trained and evaluated on small datasets with limited application…

Sound · Computer Science 2023-09-08 Zhaohui Yin , Jingguang Tian , Xinhui Hu , Xinkang Xu , Yang Xiang

Deepfakes represent a growing concern across domains such as disinformation, fraud, and non-consensual media. In particular, the rise of video conference and identity-driven attacks in high-stakes scenarios--such as impostor hiring--demands…

Computer Vision and Pattern Recognition · Computer Science 2026-04-08 Sarah Barrington , Maty Bohacek , Hany Farid

The two popular datasets ScanRefer [16] and ReferIt3D [3] connect natural language to real-world 3D data. In this paper, we curate a large-scale and complementary dataset extending both the aforementioned ones by associating all objects…

Computer Vision and Pattern Recognition · Computer Science 2023-04-04 Ahmed Abdelreheem , Kyle Olszewski , Hsin-Ying Lee , Peter Wonka , Panos Achlioptas

Inspired by humans comprehending speech in a multi-modal manner, various audio-visual datasets have been constructed. However, most existing datasets focus on English, induce dependencies with various prediction models during dataset…

One of the factors that have hindered progress in the areas of sign language recognition, translation, and production is the absence of large annotated datasets. Towards this end, we introduce How2Sign, a multimodal and multiview continuous…

Computer Vision and Pattern Recognition · Computer Science 2021-04-02 Amanda Duarte , Shruti Palaskar , Lucas Ventura , Deepti Ghadiyaram , Kenneth DeHaan , Florian Metze , Jordi Torres , Xavier Giro-i-Nieto

This paper presents a database of human faces for persons wearing spectacles. The database consists of images of faces having significant variations with respect to illumination, head pose, skin color, facial expressions and sizes, and…

Computer Vision and Pattern Recognition · Computer Science 2016-08-12 Anirban Dasgupta , Shubhobrata Bhattacharya , Aurobinda Routray

Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text descriptions for visual frames, considering audio to be…

This paper presents a novel metric learning approach to address the performance gap between normal and silent speech in visual speech recognition (VSR). The difference in lip movements between the two poses a challenge for existing VSR…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-17 Sara Kashiwagi , Keitaro Tanaka , Qi Feng , Shigeo Morishima

Violence detection has been studied in computer vision for years. However, previous work are either superficial, e.g., classification of short-clips, and the single scenario, or undersupplied, e.g., the single modality, and hand-crafted…

Computer Vision and Pattern Recognition · Computer Science 2020-07-14 Peng Wu , Jing Liu , Yujia Shi , Yujia Sun , Fangtao Shao , Zhaoyang Wu , Zhiwei Yang
‹ Prev 1 8 9 10 Next ›