中文
相关论文

相关论文: Cross-Utterance Conditioned VAE for Speech Generat…

200 篇论文

Deep generative models have been enjoying success in modeling continuous data. However it remains challenging to capture the representations for discrete structures with formal grammars and semantics, e.g., computer programs and molecular…

机器学习 · 计算机科学 2018-02-27 Hanjun Dai , Yingtao Tian , Bo Dai , Steven Skiena , Le Song

Recently, a complex variational autoencoder (VAE)-based single-channel speech enhancement system based on the DCCRN architecture has been proposed. In this system, a noise suppression VAE (NSVAE) learns to extract clean speech…

音频与语音处理 · 电气工程与系统科学 2026-02-03 Jiatong Li , Simon Doclo

Speech enhancement plays an essential role in various applications, and the integration of visual information has been demonstrated to bring substantial advantages. However, the majority of current research concentrates on the examination…

声音 · 计算机科学 2025-04-03 Xinyuan Qian , Jiaran Gao , Yaodan Zhang , Qiquan Zhang , Hexin Liu , Leibny Paola Garcia , Haizhou Li

Speech enhancement aims to improve the quality of speech signals in terms of quality and intelligibility, and speech editing refers to the process of editing the speech according to specific user needs. In this paper, we propose a Unified…

声音 · 计算机科学 2023-10-03 Muqiao Yang , Chunlei Zhang , Yong Xu , Zhongweiyang Xu , Heming Wang , Bhiksha Raj , Dong Yu

Recent research has delved into speech enhancement (SE) approaches that leverage audio embeddings from pre-trained models, diverging from time-frequency masking or signal prediction techniques. This paper introduces an efficient and…

音频与语音处理 · 电气工程与系统科学 2025-06-16 Xingwei Sun , Heinrich Dinkel , Yadong Niu , Linzhang Wang , Junbo Zhang , Jian Luan

Text-to-Speech (TTS) synthesis plays an important role in human-computer interaction. Currently, most TTS technologies focus on the naturalness of speech, namely,making the speeches sound like humans. However, the key tasks of the…

声音 · 计算机科学 2021-05-11 Jinyin Chen , Linhui Ye , Zhaoyan Ming

Emotional voice conversion (EVC) aims to convert the emotion of speech from one state to another while preserving the linguistic content and speaker identity. In this paper, we study the disentanglement and recomposition of emotional…

声音 · 计算机科学 2020-11-05 Kun Zhou , Berrak Sisman , Haizhou Li

Although high-fidelity speech can be obtained for intralingual speech synthesis, cross-lingual text-to-speech (CTTS) is still far from satisfactory as it is difficult to accurately retain the speaker timbres(i.e. speaker similarity) and…

声音 · 计算机科学 2023-06-27 Sen Liu , Yiwei Guo , Chenpeng Du , Xie Chen , Kai Yu

Human speech can be characterized by different components, including semantic content, speaker identity and prosodic information. Significant progress has been made in disentangling representations for semantic content and speaker identity…

声音 · 计算机科学 2023-09-27 Leyuan Qu , Taihao Li , Cornelius Weber , Theresa Pekarek-Rosin , Fuji Ren , Stefan Wermter

This chapter presents a novel approach to brain-to-speech (BTS) synthesis from intracranial electroencephalography (iEEG) data, emphasizing prosody-aware feature engineering and advanced transformer-based models for high-fidelity speech…

信号处理 · 电气工程与系统科学 2026-04-08 Mohammed Salah Al-Radhi , Géza Németh , Andon Tchechmedjiev , Binbin Xu

Cross-speaker style transfer is crucial to the applications of multi-style and expressive speech synthesis at scale. It does not require the target speakers to be experts in expressing all styles and to collect corresponding recordings for…

声音 · 计算机科学 2021-07-28 Shifeng Pan , Lei He

Current voice conversion (VC) methods can successfully convert timbre of the audio. As modeling source audio's prosody effectively is a challenging task, there are still limitations of transferring source style to the converted speech. This…

音频与语音处理 · 电气工程与系统科学 2021-06-29 Zhichao Wang , Xinyong Zhou , Fengyu Yang , Tao Li , Hongqiang Du , Lei Xie , Wendong Gan , Haitao Chen , Hai Li

Using a text description as prompt to guide the generation of text or images (e.g., GPT-3 or DALLE-2) has drawn wide attention recently. Beyond text and image generation, in this work, we explore the possibility of utilizing text…

音频与语音处理 · 电气工程与系统科学 2022-11-23 Zhifang Guo , Yichong Leng , Yihan Wu , Sheng Zhao , Xu Tan

Text to speech (TTS) has made rapid progress in both academia and industry in recent years. Some questions naturally arise that whether a TTS system can achieve human-level quality, how to define/judge that quality and how to achieve it. In…

音频与语音处理 · 电气工程与系统科学 2022-05-11 Xu Tan , Jiawei Chen , Haohe Liu , Jian Cong , Chen Zhang , Yanqing Liu , Xi Wang , Yichong Leng , Yuanhao Yi , Lei He , Frank Soong , Tao Qin , Sheng Zhao , Tie-Yan Liu

With rapid globalization, the need to build inclusive and representative speech technology cannot be overstated. Accent is an important aspect of speech that needs to be taken into consideration while building inclusive speech synthesizers.…

音频与语音处理 · 电气工程与系统科学 2024-10-01 Jan Melechovsky , Ambuj Mehrish , Berrak Sisman , Dorien Herremans

Many speech processing tasks involve measuring the acoustic similarity between speech segments. Acoustic word embeddings (AWE) allow for efficient comparisons by mapping speech segments of arbitrary duration to fixed-dimensional vectors.…

计算与语言 · 计算机科学 2020-12-15 Lisa van Staden , Herman Kamper

We present a neural text-to-speech (TTS) method that models natural vocal effort variation to improve the intelligibility of synthetic speech in the presence of noise. The method consists of first measuring the spectral tilt of unlabeled…

音频与语音处理 · 电气工程与系统科学 2022-03-30 Tuomo Raitio , Petko Petkov , Jiangchuan Li , Muhammed Shifas , Andrea Davis , Yannis Stylianou

Emotional voice conversion (EVC) is one way to generate expressive synthetic speech. Previous approaches mainly focused on modeling one-to-one mapping, i.e., conversion from one emotional state to another emotional state, with Mel-cepstral…

音频与语音处理 · 电气工程与系统科学 2020-04-09 Songxiang Liu , Yuewen Cao , Helen Meng

We introduce HybridVC, a voice conversion (VC) framework built upon a pre-trained conditional variational autoencoder (CVAE) that combines the strengths of a latent model with contrastive learning. HybridVC supports text and audio prompts,…

声音 · 计算机科学 2024-09-26 Xinlei Niu , Jing Zhang , Charles Patrick Martin

Generating expressive and contextually appropriate prosody remains a challenge for modern text-to-speech (TTS) systems. This is particularly evident for long, multi-sentence inputs. In this paper, we examine simple extensions to a…

音频与语音处理 · 电气工程与系统科学 2022-06-30 Peter Makarov , Ammar Abbas , Mateusz Łajszczak , Arnaud Joly , Sri Karlapati , Alexis Moinet , Thomas Drugman , Penny Karanasou