中文
相关论文

相关论文: Content-Dependent Fine-Grained Speaker Embedding f…

200 篇论文

We present a meta-learning approach for adaptive text-to-speech (TTS) with few data. During training, we learn a multi-speaker model using a shared conditional WaveNet core and independent learned embeddings for each speaker. The aim of…

This paper proposes a zero-shot learning approach for audio classification based on the textual information about class labels without any audio samples from target classes. We propose an audio classification system built on the bilinear…

机器学习 · 计算机科学 2019-08-08 Huang Xie , Tuomas Virtanen

Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal prediction techniques. This paper introduces an efficient and…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Xingwei Sun , Heinrich Dinkel , Yadong Niu , Linzhang Wang , Junbo Zhang , Jian Luan

Language models can be viewed as functions that embed text into Euclidean space, where the quality of the embedding vectors directly determines model performance, training such neural networks involves various uncertainties. This paper…

计算与语言 · 计算机科学 2025-03-31 Yifei Duan , Raphael Shang , Deng Liang , Yongqiang Cai

While the use of deep neural networks has significantly boosted speaker recognition performance, it is still challenging to separate speakers in poor acoustic environments. Here speech enhancement methods have traditionally allowed improved…

音频与语音处理 · 电气工程与系统科学 2020-08-28 Yanpei Shi , Qiang Huang , Thomas Hain

Though significant progress has been made for the voice conversion (VC) of typical speech, VC for atypical speech, e.g., dysarthric and second-language (L2) speech, remains a challenge, since it involves correcting for atypical prosody…

音频与语音处理 · 电气工程与系统科学 2021-07-26 Disong Wang , Songxiang Liu , Lifa Sun , Xixin Wu , Xunying Liu , Helen Meng

Speaker modeling is essential for many related tasks, such as speaker recognition and speaker diarization. The dominant modeling approach is fixed-dimensional vector representation, i.e., speaker embedding. This paper introduces a research…

In this paper, we propose a method of speaker adaption with intuitive prosodic features for statistical parametric speech synthesis. The intuitive prosodic features employed in this method include pitch, pitch range, speech rate and energy…

声音 · 计算机科学 2022-03-03 Pengyu Cheng , Zhenhua Ling

In recent years, several text-to-speech systems have been proposed to synthesize natural speech in zero-shot, few-shot, and low-resource scenarios. However, these methods typically require training with data from many different speakers.…

音频与语音处理 · 电气工程与系统科学 2025-06-03 Kishor Kayyar Lakshminarayana , Frank Zalkow , Christian Dittmar , Nicola Pia , Emanuel A. P. Habets

This paper proposes joint speaker feature learning methods for zero-shot adaptation of audio-visual multichannel speech separation and recognition systems. xVector and ECAPA-TDNN speaker encoders are connected using purpose-built fusion…

The objective of this paper is to learn representations of speaker identity without access to manually annotated data. To do so, we develop a self-supervised learning objective that exploits the natural cross-modal synchrony between faces…

音频与语音处理 · 电气工程与系统科学 2020-05-05 Arsha Nagrani , Joon Son Chung , Samuel Albanie , Andrew Zisserman

We introduce a new automatic evaluation method for speaker similarity assessment, that is consistent with human perceptual scores. Modern neural text-to-speech models require a vast amount of clean training data, which is why many solutions…

声音 · 计算机科学 2022-07-04 Deja Kamil , Sanchez Ariadna , Roth Julian , Cotescu Marius

Children speech recognition is indispensable but challenging due to the diversity of children's speech. In this paper, we propose a filter-based discriminative autoencoder for acoustic modeling. To filter out the influence of various…

计算与语言 · 计算机科学 2022-05-24 Chiang-Lin Tai , Hung-Shin Lee , Yu Tsao , Hsin-Min Wang

Transformer models have been used in automatic speech recognition (ASR) successfully and yields state-of-the-art results. However, its performance is still affected by speaker mismatch between training and test data. Further finetuning a…

音频与语音处理 · 电气工程与系统科学 2021-10-19 Yingzhu Zhao , Chongjia Ni , Cheung-Chi Leung , Shafiq Joty , Eng Siong Chng , Bin Ma

Recent studies on pronunciation scoring have explored the effect of introducing phone embeddings as reference pronunciation, but mostly in an implicit manner, i.e., addition or concatenation of reference phone embedding and actual…

音频与语音处理 · 电气工程与系统科学 2023-03-14 Wei Liu , Kaiqi Fu , Xiaohai Tian , Shuju Shi , Wei Li , Zejun Ma , Tan Lee

Speaker-aware source separation methods are promising workarounds for major difficulties such as arbitrary source permutation and unknown number of sources. However, it remains challenging to achieve satisfying performance provided a very…

声音 · 计算机科学 2018-07-25 Jun Wang , Jie Chen , Dan Su , Lianwu Chen , Meng Yu , Yanmin Qian , Dong Yu

Dominant researches adopt supervised training for speaker extraction, while the scarcity of ideally clean corpus and channel mismatch problem are rarely considered. To this end, we propose speaker-aware mixture of mixtures training (SAMoM),…

音频与语音处理 · 电气工程与系统科学 2022-04-18 Zifeng Zhao , Rongzhi Gu , Dongchao Yang , Jinchuan Tian , Yuexian Zou

Expressive zero-shot voice conversion (VC) is a critical and challenging task that aims to transform the source timbre into an arbitrary unseen speaker while preserving the original content and expressive qualities. Despite recent progress…

声音 · 计算机科学 2025-01-13 Yuguang Yang , Yu Pan , Jixun Yao , Xiang Zhang , Jianhao Ye , Hongbin Zhou , Lei Xie , Lei Ma , Jianjun Zhao

In this paper we propose a new method of speaker diarization that employs a deep learning architecture to learn speaker embeddings. In contrast to the traditional approaches that build their speaker embeddings using manually hand-crafted…

声音 · 计算机科学 2017-09-18 Pawel Cyrta , Tomasz Trzciński , Wojciech Stokowiec

The task of video-to-speech aims to translate silent video of lip movement to its corresponding audio signal. Previous approaches to this task are generally limited to the case of a single speaker, but a method that accounts for multiple…

音频与语音处理 · 电气工程与系统科学 2021-05-21 Dan Oneata , Adriana Stan , Horia Cucu