English
Related papers

Related papers: Speaker Recognition in Realistic Scenario Using Mu…

200 papers

The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective that exploits the natural cross-modal synchrony between faces…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-05 Arsha Nagrani , Joon Son Chung , Samuel Albanie , Andrew Zisserman

Emotion recognition has become an important field of research in the human-computer interactions domain. The latest advancements in the field show that combining visual with audio information lead to better results if compared to the case…

Computer Vision and Pattern Recognition · Computer Science 2020-03-03 Nicolae-Catalin Ristea , Liviu Cristian Dutu , Anamaria Radoi

We learn rich natural sound representations by capitalizing on large amounts of unlabeled sound data collected in the wild. We leverage the natural synchronization between vision and sound to learn an acoustic representation using…

Computer Vision and Pattern Recognition · Computer Science 2016-10-31 Yusuf Aytar , Carl Vondrick , Antonio Torralba

Although many previous studies have carried out multimodal learning with real-time MRI data that captures the audio-visual kinematics of the vocal tract during speech, these studies have been limited by their reliance on multi-speaker…

Machine Learning · Computer Science 2025-06-02 Sean Foley , Hong Nguyen , Jihwan Lee , Sudarsana Reddy Kadiri , Dani Byrd , Louis Goldstein , Shrikanth Narayanan

Self-supervised learning has become increasingly important to leverage the abundance of unlabeled data available on platforms like YouTube. Whereas most existing approaches learn low-level representations, we propose a joint…

Computer Vision and Pattern Recognition · Computer Science 2019-09-13 Chen Sun , Austin Myers , Carl Vondrick , Kevin Murphy , Cordelia Schmid

In this paper, a novel Convolutional Neural Network architecture has been developed for speaker verification in order to simultaneously capture and discard speaker and non-speaker information, respectively. In training phase, the network is…

Audio and Speech Processing · Electrical Eng. & Systems 2018-08-13 Hossein Salehghaffari

Advances in deep learning have resulted in state-of-the-art performance for many audio classification tasks but, unlike humans, these systems traditionally require large amounts of data to make accurate predictions. Not every person or…

Audio and Speech Processing · Electrical Eng. & Systems 2020-12-04 Piper Wolters , Chris Careaga , Brian Hutchinson , Lauren Phillips

VoxCeleb datasets are widely used in speaker recognition studies. Our work serves two purposes. First, we provide speaker age labels and (an alternative) annotation of speaker gender. Second, we demonstrate the use of this metadata by…

Machine Learning · Computer Science 2021-12-21 Khaled Hechmi , Trung Ngo Trong , Ville Hautamaki , Tomi Kinnunen

Neural network-based dialog systems are attracting increasing attention in both academia and industry. Recently, researchers have begun to realize the importance of speaker modeling in neural dialog systems, but there lacks established…

Computation and Language · Computer Science 2018-10-01 Zhao Meng , Lili Mou , Zhi Jin

Recent advances in deep learning have facilitated the design of speaker verification systems that directly input raw waveforms. For example, RawNet extracts speaker embeddings from raw waveforms, which simplifies the process pipeline and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-05-08 Jee-weon Jung , Seung-bin Kim , Hye-jin Shim , Ju-ho Kim , Ha-Jin Yu

Recent advances in unsupervised speech representation learning discover new approaches and provide new state-of-the-art for diverse types of speech processing tasks. This paper presents an investigation of using wav2vec 2.0 deep speech…

Recent success of the Tacotron speech synthesis architecture and its variants in producing natural sounding multi-speaker synthesized speech has raised the exciting possibility of replacing expensive, manually transcribed, domain-specific,…

Computation and Language · Computer Science 2019-09-27 Andrew Rosenberg , Yu Zhang , Bhuvana Ramabhadran , Ye Jia , Pedro Moreno , Yonghui Wu , Zelin Wu

In this paper, we introduce a pretrained audio-visual Transformer trained on more than 500k utterances from nearly 4000 celebrities from the VoxCeleb2 dataset for human behavior understanding. The model aims to capture and extract useful…

Multimedia · Computer Science 2022-01-25 Minh Tran , Mohammad Soleymani

This paper summarizes the applied deep learning practices in the field of speaker recognition, both verification and identification. Speaker recognition has been a widely used field topic of speech technology. Many research works have been…

Audio and Speech Processing · Electrical Eng. & Systems 2022-09-27 Dávid Sztahó , György Szaszák , András Beke

Speaker recognition technology is applied to various tasks, from personal virtual assistants to secure access systems. However, the robustness of these systems against adversarial attacks, particularly to additive perturbations, remains a…

Sound · Computer Science 2024-12-19 Dmitrii Korzh , Elvir Karimov , Mikhail Pautov , Oleg Y. Rogov , Ivan Oseledets

Recently, there has been growing interest in multi-speaker speech recognition, where the utterances of multiple speakers are recognized from their mixture. Promising techniques have been proposed for this task, but earlier works have…

Sound · Computer Science 2018-05-16 Hiroshi Seki , Takaaki Hori , Shinji Watanabe , Jonathan Le Roux , John R. Hershey

The deep learning based time-domain models, e.g. Conv-TasNet, have shown great potential in both single-channel and multi-channel speech enhancement. However, many experiments on the time-domain speech enhancement model are done in…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-28 Wangyou Zhang , Jing Shi , Chenda Li , Shinji Watanabe , Yanmin Qian

Recently, researchers have gradually realized that in some cases, the self-supervised pre-training on large-scale Internet data is better than that of high-quality/manually labeled data sets, and multimodal/large models are better than…

Sound · Computer Science 2023-08-08 Sen Fang , Yangjian Wu , Bowen Gao , Jingwen Cai , Teik Toe Teoh

This paper introduces a new multi-modal dataset for visual and audio-visual speech recognition. It includes face tracks from over 400 hours of TED and TEDx videos, along with the corresponding subtitles and word alignment boundaries. The…

Computer Vision and Pattern Recognition · Computer Science 2018-10-30 Triantafyllos Afouras , Joon Son Chung , Andrew Zisserman

Acoustic emotion recognition aims to categorize the affective state of the speaker and is still a difficult task for machine learning models. The difficulties come from the scarcity of training data, general subjectivity in emotion…

Computation and Language · Computer Science 2018-04-02 Egor Lakomkin , Cornelius Weber , Sven Magg , Stefan Wermter