中文
相关论文

相关论文: VoxCeleb2: Deep Speaker Recognition

200 篇论文

Deep learning is progressively gaining popularity as a viable alternative to i-vectors for speaker recognition. Promising results have been recently obtained with Convolutional Neural Networks (CNNs) when fed by raw speech samples directly.…

音频与语音处理 · 电气工程与系统科学 2019-08-12 Mirco Ravanelli , Yoshua Bengio

This paper presents VoxCeleb-ESP, a collection of pointers and timestamps to YouTube videos facilitating the creation of a novel speaker recognition dataset. VoxCeleb-ESP captures real-world scenarios, incorporating diverse speaking styles,…

This paper describes our submission to the Second Clarity Enhancement Challenge (CEC2), which consists of target speech enhancement for hearing-aid (HA) devices in noisy-reverberant environments with multiple interferers such as music and…

音频与语音处理 · 电气工程与系统科学 2023-02-17 Samuele Cornell , Zhong-Qiu Wang , Yoshiki Masuyama , Shinji Watanabe , Manuel Pariente , Nobutaka Ono

This paper tackles the problem of the heavy dependence of clean speech data required by deep learning based audio-denoising methods by showing that it is possible to train deep speech denoising networks using only noisy speech samples.…

声音 · 计算机科学 2021-09-21 Madhav Mahesh Kashyap , Anuj Tambwekar , Krishnamoorthy Manohara , S Natarajan

Current state-of-the-art speech recognition models are trained to map acoustic signals into sub-lexical units. While these models demonstrate superior performance, they remain vulnerable to out-of-distribution conditions such as background…

声音 · 计算机科学 2024-10-10 Sagarika Alavilli , Annesya Banerjee , Gasser Elbanna , Annika Magaro

Today's Automatic Speech Recognition systems only rely on acoustic signals and often don't perform well under noisy conditions. Performing multi-modal speech recognition - processing acoustic speech signals and lip-reading video…

计算机视觉与模式识别 · 计算机科学 2018-03-14 Matthijs Van keirsbilck , Bert Moons , Marian Verhelst

A deep learning approach has been proposed recently to derive speaker identifies (d-vector) by a deep neural network (DNN). This approach has been applied to text-dependent speaker recognition tasks and shows reasonable performance gains…

计算与语言 · 计算机科学 2015-06-30 Lantian Li , Yiye Lin , Zhiyong Zhang , Dong Wang

We introduce a seemingly impossible task: given only an audio clip of someone speaking, decide which of two face images is the speaker. In this paper we study this, and a number of related cross-modal tasks, aimed at answering the question:…

计算机视觉与模式识别 · 计算机科学 2018-04-04 Arsha Nagrani , Samuel Albanie , Andrew Zisserman

Recently, fine-tuning large pre-trained Transformer models using downstream datasets has received a rising interest. Despite their success, it is still challenging to disentangle the benefits of large-scale datasets and Transformer…

音频与语音处理 · 电气工程与系统科学 2023-05-19 Junyi Peng , Oldřich Plchot , Themos Stafylakis , Ladislav Mošner , Lukáš Burget , Jan Černocký

We investigate robustness properties of pre-trained neural models for automatic speech recognition. Real life data in machine learning is usually very noisy and almost never clean, which can be attributed to various factors depending on the…

计算与语言 · 计算机科学 2022-08-19 Goutham Rajendran , Wei Zou

The availability of large labeled datasets has allowed Convolutional Network models to achieve impressive recognition results. However, in many settings manual annotation of the data is impractical; instead our data has noisy labels, i.e.…

计算机视觉与模式识别 · 计算机科学 2015-04-13 Sainbayar Sukhbaatar , Joan Bruna , Manohar Paluri , Lubomir Bourdev , Rob Fergus

Speech audio in the wild is often processed by post-production effects, but existing speech datasets rarely provide precise annotations of effects and parameters, limiting systematic study. We introduce VoxEffects, a speech audio effects…

音频与语音处理 · 电气工程与系统科学 2026-04-15 Zhe Zhang , Yigitcan Özer , Junichi Yamagishi

The success of automatic speaker verification shows that discriminative speaker representations can be extracted from neutral speech. However, as a kind of non-verbal voice, laughter should also carry speaker information intuitively. Thus,…

音频与语音处理 · 电气工程与系统科学 2023-11-21 Yuke Lin , Xiaoyi Qin , Huahua Cui , Zhenyi Zhu , Ming Li

Todays interactive devices such as smart-phone assistants and smart speakers often deal with short-duration speech segments. As a result, speaker recognition systems integrated into such devices will be much better suited with models…

音频与语音处理 · 电气工程与系统科学 2019-07-25 Amirhossein Hajavi , Ali Etemad

We propose a novel deep training algorithm for joint representation of audio and visual information which consists of a single stream network (SSNet) coupled with a novel loss function to learn a shared deep latent space representation of…

计算机视觉与模式识别 · 计算机科学 2019-09-20 Shah Nawaz , Muhammad Kamran Janjua , Ignazio Gallo , Arif Mahmood , Alessandro Calefati

Accurately detecting voiced intervals in speech signals is a critical step in pitch tracking and has numerous applications. While conventional signal processing methods and deep learning algorithms have been proposed for this task, their…

音频与语音处理 · 电气工程与系统科学 2023-12-07 Yixuan Zhang , Heming Wang , DeLiang Wang

In this paper, a novel method using 3D Convolutional Neural Network (3D-CNN) architecture has been proposed for speaker verification in the text-independent setting. One of the main challenges is the creation of the speaker models. Most of…

计算机视觉与模式识别 · 计算机科学 2018-06-08 Amirsina Torfi , Jeremy Dawson , Nasser M. Nasrabadi

Deep neural networks can learn complex and abstract representations, that are progressively obtained by combining simpler ones. A recent trend in speech and speaker recognition consists in discovering these representations starting from raw…

音频与语音处理 · 电气工程与系统科学 2019-02-26 Mirco Ravanelli , Yoshua Bengio

Recently, neural vocoders have been widely used in speech synthesis tasks, including text-to-speech and voice conversion. However, when encountering data distribution mismatch between training and inference, neural vocoders trained on real…

声音 · 计算机科学 2020-08-21 Po-chun Hsu , Chun-hsuan Wang , Andy T. Liu , Hung-yi Lee

Text-to-speech models trained on large-scale datasets have demonstrated impressive in-context learning capabilities and naturalness. However, control of speaker identity and style in these models typically requires conditioning on reference…

声音 · 计算机科学 2024-02-08 Dan Lyth , Simon King