中文
相关论文

相关论文: Multi-Target Emotional Voice Conversion With Neura…

200 篇论文

Recently the state-of-the-art text-to-speech synthesis systems have shifted to a two-model approach: a sequence-to-sequence model to predict a representation of speech (typically mel-spectrograms), followed by a 'neural vocoder' model which…

音频与语音处理 · 电气工程与系统科学 2020-12-18 Jonas Rohnke , Tom Merritt , Jaime Lorenzo-Trueba , Adam Gabrys , Vatsal Aggarwal , Alexis Moinet , Roberto Barra-Chicote

State-of-the-art speech synthesis models try to get as close as possible to the human voice. Hence, modelling emotions is an essential part of Text-To-Speech (TTS) research. In our work, we selected FastSpeech2 as the starting point and…

音频与语音处理 · 电气工程与系统科学 2023-07-04 Daria Diatlova , Vitaly Shutov

Recent advances in text-to-speech (TTS) have yielded remarkable improvements in naturalness and intelligibility. Building on these achievements, research has increasingly shifted toward enhancing the expressiveness of generated speech, such…

声音 · 计算机科学 2025-12-23 Pengchao Feng , Yao Xiao , Ziyang Ma , Zhikang Niu , Shuai Fan , Yao Li , Sheng Wang , Xie Chen

Audio-visual speech enhancement aims to extract clean speech from a noisy environment by leveraging not only the audio itself but also the target speaker's lip movements. This approach has been shown to yield improvements over audio-only…

Emotional Text-to-Speech (E-TTS) synthesis has garnered significant attention in recent years due to its potential to revolutionize human-computer interaction. However, current E-TTS approaches often struggle to capture the intricacies of…

计算与语言 · 计算机科学 2025-02-20 Zhi-Qi Cheng , Xiang Li , Jun-Yan He , Junyao Chen , Xiaomao Fan , Xiaojiang Peng , Alexander G. Hauptmann

Emotional talking head synthesis aims to generate talking portrait videos with vivid expressions. Existing methods still exhibit limitations in control flexibility, motion naturalness, and expression quality. Moreover, currently available…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Yiguo Jiang , Xiaodong Cun , Yong Zhang , Yudian Zheng , Fan Tang , Chi-Man Pun

This paper presents an emotion-regularized conditional variational autoencoder (Emo-CVAE) model for generating emotional conversation responses. In conventional CVAE-based emotional response generation, emotion labels are simply used as…

计算与语言 · 计算机科学 2021-04-20 Yu-Ping Ruan , Zhen-Hua Ling

Affect is an emotional characteristic encompassing valence, arousal, and intensity, and is a crucial attribute for enabling authentic conversations. While existing text-to-speech (TTS) and speech-to-speech systems rely on strength embedding…

We propose a linear prediction (LP)-based waveform generation method via WaveNet vocoding framework. A WaveNet-based neural vocoder has significantly improved the quality of parametric text-to-speech (TTS) systems. However, it is…

音频与语音处理 · 电气工程与系统科学 2020-03-05 Min-Jae Hwang , Frank Soong , Eunwoo Song , Xi Wang , Hyeonjoo Kang , Hong-Goo Kang

In recent years, prompting has quickly become one of the standard ways of steering the outputs of generative machine learning models, due to its intuitive use of natural language. In this work, we propose a system conditioned on embeddings…

计算与语言 · 计算机科学 2024-06-13 Thomas Bott , Florian Lux , Ngoc Thang Vu

We present an Audio-Visual Language Model (AVLM) for expressive speech generation by integrating full-face visual cues into a pre-trained expressive speech model. We explore multiple visual encoders and multimodal fusion strategies during…

计算与语言 · 计算机科学 2025-08-29 Weiting Tan , Jiachen Lian , Hirofumi Inaguma , Paden Tomasello , Philipp Koehn , Xutai Ma

Zero-shot voice conversion (VC) aims to transfer the source speaker timbre to arbitrary unseen target speaker timbre, while keeping the linguistic content unchanged. Although the voice of generated speech can be controlled by providing the…

声音 · 计算机科学 2024-01-31 Junjie Li , Yiwei Guo , Xie Chen , Kai Yu

Emotion recognition in conversation (ERC) aims to analyze the speaker's state and identify their emotion in the conversation. Recent works in ERC focus on context modeling but ignore the representation of contextual emotional tendency. In…

计算与语言 · 计算机科学 2022-03-28 Zaijing Li , Fengxiao Tang , Ming Zhao , Yusen Zhu

Language models (LMs) have recently flourished in natural language processing and computer vision, generating high-fidelity texts or images in various tasks. In contrast, the current speech generative models are still struggling regarding…

声音 · 计算机科学 2023-10-13 Xinfa Zhu , Yuanjun Lv , Yi Lei , Tao Li , Wendi He , Hongbin Zhou , Heng Lu , Lei Xie

In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional…

音频与语音处理 · 电气工程与系统科学 2019-07-08 Noé Tits

Recent advances in neural autoregressive models have improve the performance of speech synthesis (SS). However, as they lack the ability to model global characteristics of speech (such as speaker individualities or speaking styles),…

计算与语言 · 计算机科学 2019-02-12 Kei Akuzawa , Yusuke Iwasawa , Yutaka Matsuo

This paper proposes speaker-adaptive neural vocoders for parametric text-to-speech (TTS) systems. Recently proposed WaveNet-based neural vocoding systems successfully generate a time sequence of speech signal with an autoregressive…

音频与语音处理 · 电气工程与系统科学 2020-08-04 Eunwoo Song , Jin-Seob Kim , Kyungguen Byun , Hong-Goo Kang

Variational auto-encoders (VAEs) are deep generative latent variable models that can be used for learning the distribution of complex data. VAEs have been successfully used to learn a probabilistic prior over speech signals, which is then…

声音 · 计算机科学 2020-12-18 Mostafa Sadeghi , Simon Leglaive , Xavier Alameda-PIneda , Laurent Girin , Radu Horaud

We present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech synthesis. The system comprises five major building blocks:…

Voice Conversion research in recent times has increasingly focused on improving the zero-shot capabilities of existing methods. Despite remarkable advancements, current architectures still tend to struggle in zero-shot cross-lingual…

声音 · 计算机科学 2025-05-26 Advait Joglekar , Divyanshu Singh , Rooshil Rohit Bhatia , S. Umesh