中文
相关论文

相关论文: Multi-Speaker Expressive Speech Synthesis via Mult…

200 篇论文

The end-to-end speech synthesis model can directly take an utterance as reference audio, and generate speech from the text with prosody and speaker characteristics similar to the reference audio. However, an appropriate acoustic embedding…

声音 · 计算机科学 2021-10-12 Cheng Gong , Longbiao Wang , Zhenhua Ling , Ju Zhang , Jianwu Dang

Speech signals encompass various information across multiple levels including content, speaker, and style. Disentanglement of these information, although challenging, is important for applications such as voice conversion. The contrastive…

音频与语音处理 · 电气工程与系统科学 2024-09-06 Yuying Xie , Michael Kuhlmann , Frederik Rautenberg , Zheng-Hua Tan , Reinhold Haeb-Umbach

Emotional voice conversion (EVC) is one way to generate expressive synthetic speech. Previous approaches mainly focused on modeling one-to-one mapping, i.e., conversion from one emotional state to another emotional state, with Mel-cepstral…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Songxiang Liu , Yuewen Cao , Helen Meng

Target speaker extraction (TSE) aims to isolate a specific speaker's voice from multi-speaker mixtures. Despite strong benchmark results, real-world performance often degrades due to different interacting factors. Previous curriculum…

声音 · 计算机科学 2026-03-06 Yun Liu , Xuechen Liu , Xiaoxiao Miao , Junichi Yamagishi

Emotional voice conversion (EVC) aims to change the emotional state of an utterance while preserving the linguistic content and speaker identity. In this paper, we propose a novel 2-stage training strategy for sequence-to-sequence emotional…

计算与语言 · 计算机科学 2021-06-10 Kun Zhou , Berrak Sisman , Haizhou Li

Target speaker extraction (TSE) relies on a reference cue of the target to extract the target speech from a speech mixture. While a speaker embedding is commonly used as the reference cue, such embedding pre-trained with a large number of…

音频与语音处理 · 电气工程与系统科学 2024-12-12 Ke Zhang , Junjie Li , Shuai Wang , Yangjie Wei , Yi Wang , Yannan Wang , Haizhou Li

Given a pair of source and reference speech recordings, speech-to-speech (S2S) emotion style transfer involves the generation of an output speech that mimics the emotion characteristics of the reference while preserving the content and…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Soumya Dutta , Avni Jain , Sriram Ganapathy

The goal of this work is to generate natural speech in multiple languages while maintaining the same speaker identity, a task known as cross-lingual speech synthesis. A key challenge of cross-lingual speech synthesis is the language-speaker…

音频与语音处理 · 电气工程与系统科学 2024-12-31 Ji-Hoon Kim , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

Overlapping speech diarization is always treated as a multi-label classification problem. In this paper, we reformulate this task as a single-label prediction problem by encoding the multi-speaker labels with power set. Specifically, we…

声音 · 计算机科学 2021-11-30 Zhihao Du , Shiliang Zhang , Siqi Zheng , Weilong Huang , Ming Lei

Recent advancements in Text-to-Speech (TTS) systems have enabled the generation of natural and expressive speech from textual input. Accented TTS aims to enhance user experience by making the synthesized speech more relatable to minority…

音频与语音处理 · 电气工程与系统科学 2024-10-18 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

We introduce an approach to multilingual speech synthesis which uses the meta-learning concept of contextual parameter generation and produces natural-sounding multilingual speech using more languages and less training data than previous…

音频与语音处理 · 电气工程与系统科学 2020-08-04 Tomáš Nekvinda , Ondřej Dušek

In expressive speech synthesis, there are high requirements for emotion interpretation. However, it is time-consuming to acquire emotional audio corpus for arbitrary speakers due to their deduction ability. In response to this problem, this…

音频与语音处理 · 电气工程与系统科学 2021-10-12 Pengfei Wu , Junjie Pan , Chenchang Xu , Junhui Zhang , Lin Wu , Xiang Yin , Zejun Ma

This paper proposes an interesting voice and accent joint conversion approach, which can convert an arbitrary source speaker's voice to a target speaker with non-native accent. This problem is challenging as each target speaker only has…

声音 · 计算机科学 2020-11-18 Zhichao Wang , Wenshuo Ge , Xiong Wang , Shan Yang , Wendong Gan , Haitao Chen , Hai Li , Lei Xie , Xiulin Li

Previous multilingual text-to-speech (TTS) approaches have considered leveraging monolingual speaker data to enable cross-lingual speech synthesis. However, such data-efficient approaches have ignored synthesizing emotional aspects of…

音频与语音处理 · 电气工程与系统科学 2023-08-01 Xinfa Zhu , Yi Lei , Tao Li , Yongmao Zhang , Hongbin Zhou , Heng Lu , Lei Xie

Cross-lingual speech emotion recognition (SER) remains a challenging task due to differences in phonetic variability and speaker-specific expressive styles across languages. Effectively capturing emotion under such diverse conditions…

计算与语言 · 计算机科学 2025-09-26 Shreya G. Upadhyay , Carlos Busso , Chi-Chun Lee

Emotional speech synthesis aims to synthesize human voices with various emotional effects. The current studies are mostly focused on imitating an averaged style belonging to a specific emotion type. In this paper, we seek to generate speech…

计算与语言 · 计算机科学 2023-01-02 Kun Zhou , Berrak Sisman , Rajib Rana , B. W. Schuller , Haizhou Li

We describe a neural network-based system for text-to-speech (TTS) synthesis that is able to generate speech audio in the voice of many different speakers, including those unseen during training. Our system consists of three independently…

Speaker extraction requires a sample speech from the target speaker as the reference. However, enrolling a speaker with a long speech is not practical. We propose a speaker extraction technique, that performs in multiple stages to take full…

音频与语音处理 · 电气工程与系统科学 2021-04-05 Meng Ge , Chenglin Xu , Longbiao Wang , Eng Siong Chng , Jianwu Dang , Haizhou Li

We present a novel source separation model to decompose asingle-channel speech signal into two speech segments belonging to two different speakers. The proposed model is a neural network based on residual blocks, and uses learnt speaker…

声音 · 计算机科学 2019-06-25 Shuo Liu , Gil Keren , Björn Schuller

Speaker clustering is the task of identifying the unique speakers in a set of audio recordings (each belonging to exactly one speaker) without knowing who and how many speakers are present in the entire data, which is essential for speaker…

声音 · 计算机科学 2025-09-30 Chaohao Lin , Xu Zheng , Kaida Wu , Peihao Xiang , Ou Bai