中文
相关论文

相关论文: AVA-ActiveSpeaker: An Audio-Visual Dataset for Act…

200 篇论文

Communicating in noisy, multi-talker environments is challenging, especially for people with hearing impairments. Egocentric video data can potentially be used to identify a user's conversation partners, which could be used to inform…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Tobias Dorszewski , Søren A. Fuglsang , Jens Hjortkjær

We present the Multiview Extended Video with Activities (MEVA) dataset, a new and very-large-scale dataset for human activity recognition. Existing security datasets either focus on activity counts by aggregating public video disseminated…

计算机视觉与模式识别 · 计算机科学 2020-12-03 Kellie Corona , Katie Osterdahl , Roderic Collins , Anthony Hoogs

In this paper, we focus on the Audio-Visual Question Answering (AVQA) task, which aims to answer questions regarding different visual objects, sounds, and their associations in videos. The problem requires comprehensive multimodal…

计算机视觉与模式识别 · 计算机科学 2022-04-06 Guangyao Li , Yake Wei , Yapeng Tian , Chenliang Xu , Ji-Rong Wen , Di Hu

We present SpeakingFaces as a publicly-available large-scale multimodal dataset developed to support machine learning research in contexts that utilize a combination of thermal, visual, and audio data streams; examples include…

Most existing datasets for speaker identification contain samples obtained under quite constrained conditions, and are usually hand-annotated, hence limited in size. The goal of this paper is to generate a large scale text-independent…

声音 · 计算机科学 2020-11-05 Arsha Nagrani , Joon Son Chung , Andrew Zisserman

Voice Activity Detection (VAD) is the process of automatically determining whether a person is speaking and identifying the timing of their speech in an audiovisual data. Traditionally, this task has been tackled by processing either audio…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Andrea Appiani , Cigdem Beyan

Speaker verification (SV) provides billions of voice-enabled devices with access control, and ensures the security of voice-driven technologies. As a type of biometrics, it is necessary that SV is unbiased, with consistent and reliable…

音频与语音处理 · 电气工程与系统科学 2022-09-14 Wiebke Toussaint Hutiri , Lauriane Gorce , Aaron Yi Ding

State-of-the-art Active Speaker Detection (ASD) approaches heavily rely on audio and facial features to perform, which is not a sustainable approach in wild scenarios. Although these methods achieve good results in the standard…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Tiago Roxo , Joana C. Costa , Pedro R. M. Inácio , Hugo Proença

This paper proposes a powerful Visual Speech Recognition (VSR) method for multiple languages, especially for low-resource languages that have a limited number of labeled data. Different from previous methods that tried to improve the VSR…

计算机视觉与模式识别 · 计算机科学 2024-01-15 Jeong Hun Yeo , Minsu Kim , Shinji Watanabe , Yong Man Ro

Nowadays, the large amount of audio-visual content available has fostered the need to develop new robust automatic speaker diarization systems to analyse and characterise it. This kind of system helps to reduce the cost of doing this…

声音 · 计算机科学 2024-09-10 Victoria Mingote , Alfonso Ortega , Antonio Miguel , Eduardo Lleida

With the recent advancements in AI, Intelligent Virtual Assistants (IVA) have become a ubiquitous part of every home. Going forward, we are witnessing a confluence of vision, speech and dialog system technologies that are enabling the IVAs…

计算与语言 · 计算机科学 2018-12-21 Shachi H Kumar , Eda Okur , Saurav Sahay , Juan Jose Alvarado Leanos , Jonathan Huang , Lama Nachman

The goal of this paper is speaker diarisation of videos collected 'in the wild'. We make three key contributions. First, we propose an automatic audio-visual diarisation method for YouTube videos. Our method consists of active speaker…

声音 · 计算机科学 2021-08-17 Joon Son Chung , Jaesung Huh , Arsha Nagrani , Triantafyllos Afouras , Andrew Zisserman

We present a joint audio-visual model for isolating a single speech signal from a mixture of sounds such as other speakers and background noise. Solving this task using only audio as input is extremely challenging and does not provide an…

It is known that deep neural networks are vulnerable to adversarial attacks. Although Automatic Speaker Verification (ASV) built on top of deep neural networks exhibits robust performance in controlled scenarios, many studies confirm that…

声音 · 计算机科学 2024-01-17 Li Wang , Jiaqi Li , Yuhao Luo , Jiahao Zheng , Lei Wang , Hao Li , Ke Xu , Chengfang Fang , Jie Shi , Zhizheng Wu

Talking-head videos constitute a predominant content type in real-time communication, yet publicly available datasets for video processing research in this domain remain scarce and limited in signal fidelity. In this paper, we open-source a…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Babak Naderi , Ross Cutler

This work presents an extensive and detailed study on Audio-Visual Speech Recognition (AVSR) for five widely spoken languages: Chinese, Spanish, English, Arabic, and French. We have collected large-scale datasets for each language except…

计算与语言 · 计算机科学 2024-06-04 Sanath Narayan , Yasser Abdelaziz Dahou Djilali , Ankit Singh , Eustache Le Bihan , Hakim Hacid

This work presents a scalable solution to open-vocabulary visual speech recognition. To achieve this, we constructed the largest existing visual speech recognition dataset, consisting of pairs of text and video clips of faces speaking…

Active speaker detection and speech enhancement have become two increasingly attractive topics in audio-visual scenario understanding. According to their respective characteristics, the scheme of independently designed architecture has been…

声音 · 计算机科学 2022-07-08 Junwen Xiong , Yu Zhou , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha

Speaker diarization is one of the critical components of computational media intelligence as it enables a character-level analysis of story portrayals and media content understanding. Automated audio-based speaker diarization of…

多媒体 · 计算机科学 2022-03-31 Rahul Sharma , Shrikanth Narayanan

The training of deep learning-based multichannel speech enhancement and source localization systems relies heavily on the simulation of room impulse response and multichannel diffuse noise, due to the lack of large-scale real-recorded…

声音 · 计算机科学 2024-10-02 Bing Yang , Changsheng Quan , Yabo Wang , Pengyu Wang , Yujie Yang , Ying Fang , Nian Shao , Hui Bu , Xin Xu , Xiaofei Li