English
Related papers

Related papers: LRS3-TED: a large-scale dataset for visual speech …

200 papers

Different studies have shown the importance of visual cues throughout the speech perception process. In fact, the development of audiovisual approaches has led to advances in the field of speech technologies. However, although noticeable…

Computer Vision and Pattern Recognition · Computer Science 2023-11-22 David Gimeno-Gómez , Carlos-D. Martínez-Hinarejos

The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Soravit Changpinyo , Piyush Sharma , Nan Ding , Radu Soricut

Active speaker detection (ASD) in videos with multiple speakers is a challenging task as it requires learning effective audiovisual features and spatial-temporal correlations over long temporal windows. In this paper, we present SPELL, a…

Computer Vision and Pattern Recognition · Computer Science 2022-10-13 Kyle Min , Sourya Roy , Subarna Tripathi , Tanaya Guha , Somdeb Majumdar

We introduce Look and Tell, a multimodal dataset for studying referential communication across egocentric and exocentric perspectives. Using Meta Project Aria smart glasses and stationary cameras, we recorded synchronized gaze, speech, and…

Computer Vision and Pattern Recognition · Computer Science 2025-10-29 Anna Deichler , Jonas Beskow

This paper describes a system that generates speaker-annotated transcripts of meetings by using a microphone array and a 360-degree camera. The hallmark of the system is its ability to handle overlapped speech, which has been an unsolved…

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

Self-supervision has shown great potential for audio-visual speech recognition by vastly reducing the amount of labeled data required to build good systems. However, existing methods are either not entirely end-to-end or do not train joint…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-23 Jiachen Lian , Alexei Baevski , Wei-Ning Hsu , Michael Auli

Speech conveys not only linguistic information but also rich non-verbal vocal events such as laughing and crying. While semantic transcription is well-studied, the precise localization of non-verbal events remains a critical yet…

Computation and Language · Computer Science 2026-01-09 Chenchen Yang , Kexin Huang , Liwei Fan , Qian Tu , Botian Jiang , Dong Zhang , Linqi Yin , Shimin Li , Zhaoye Fei , Qinyuan Cheng , Xipeng Qiu

This paper presents an augmentation of MSCOCO dataset where speech is added to image and text. Speech captions are generated using text-to-speech (TTS) synthesis resulting in 616,767 spoken captions (more than 600h) paired with images.…

Computation and Language · Computer Science 2020-11-24 William Havard , Laurent Besacier , Olivier Rosec

Understanding human visual attention and saliency is an integral part of vision research. In this context, there is an ever-present need for fresh and diverse benchmark datasets, particularly for insight into special use cases like crowded…

Computer Vision and Pattern Recognition · Computer Science 2019-10-10 Memoona Tahira , Sobas Mehboob , Anis U. Rahman , Omar Arif

Humans are arguably one of the most important subjects in video streams, many real-world applications such as video summarization or video editing workflows often require the automatic search and retrieval of a person of interest. Despite…

Computer Vision and Pattern Recognition · Computer Science 2021-06-04 Juan Leon Alcazar , Long Mai , Federico Perazzi , Joon-Young Lee , Pablo Arbelaez , Bernard Ghanem , Fabian Caba Heilbron

Spoken named entity recognition (NER) aims to identify named entities from speech, playing an important role in speech processing. New named entities appear every day, however, annotating their Spoken NER data is costly. In this paper, we…

Computation and Language · Computer Science 2024-12-30 Jiawei Yu , Xiang Geng , Yuang Li , Mengxin Ren , Wei Tang , Jiahuan Li , Zhibin Lan , Min Zhang , Hao Yang , Shujian Huang , Jinsong Su

Visual-semantic embedding aims to find a shared latent space where related visual and textual instances are close to each other. Most current methods learn injective embedding functions that map an instance to a single point in the shared…

Computer Vision and Pattern Recognition · Computer Science 2019-07-18 Yale Song , Mohammad Soleymani

Video transcript summarization is a fundamental task for video understanding. Conventional approaches for transcript summarization are usually built upon the summarization data for written language such as news articles, while the domain…

Computation and Language · Computer Science 2021-07-16 Tengchao Lv , Lei Cui , Momcilo Vasilijevic , Furu Wei

Recent advances in audio-language models have demonstrated remarkable success on short, segment-level speech tasks. However, real-world applications such as meeting transcription, spoken document understanding, and conversational analysis…

Current datasets for video-based person re-identification (re-ID) do not include structural knowledge in form of human pose annotations for the persons of interest. Nonetheless, pose information is very helpful to disentangle useful feature…

Computer Vision and Pattern Recognition · Computer Science 2020-11-13 Andreas Doering , Di Chen , Shanshan Zhang , Bernt Schiele , Juergen Gall

This paper presents a new approach for end-to-end audio-visual multi-talker speech recognition. The approach, referred to here as the visual context attention model (VCAM), is important because it uses the available video information to…

Sound · Computer Science 2022-04-05 Richard Rose , Olivier Siohan

Traditional speaker diarization systems have primarily focused on constrained scenarios such as meetings and interviews, where the number of speakers is limited and acoustic conditions are relatively clean. To explore open-world speaker…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Liangbin Huang , Xiaohua Liao , Chaoqun Cui , Shijing Wang , Zhaolong Huang , Yanlong Du , Wenji Mao

We propose a method for learning from streaming visual data using a compact, constant size representation of all the data that was seen until a given moment. Specifically, we construct a 'coreset' representation of streaming data using a…

Computer Vision and Pattern Recognition · Computer Science 2015-11-20 Abhimanyu Dubey , Nikhil Naik , Dan Raviv , Rahul Sukthankar , Ramesh Raskar

The objective of this study is to generate high-quality speech from silent talking face videos, a task also known as video-to-speech synthesis. A significant challenge in video-to-speech synthesis lies in the substantial modality gap…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-24 Ji-Hoon Kim , Jeongsoo Choi , Jaehun Kim , Chaeyoung Jung , Joon Son Chung