中文
相关论文

相关论文: ZSDEVC: Zero-Shot Diffusion-based Emotional Voice …

200 篇论文

In realistic speech enhancement settings for end-user devices, we often encounter only a few speakers and noise types that tend to reoccur in the specific acoustic environment. We propose a novel personalized speech enhancement method to…

音频与语音处理 · 电气工程与系统科学 2021-05-11 Sunwoo Kim , Minje Kim

Interpreting EEG signals linked to spoken language presents a complex challenge, given the data's intricate temporal and spatial attributes, as well as the various noise factors. Denoising diffusion probabilistic models (DDPMs), which have…

计算与语言 · 计算机科学 2023-11-15 Soowon Kim , Seo-Hyun Lee , Young-Eun Lee , Ji-Won Lee , Ji-Ha Park , Seong-Whan Lee

Modern speech systems increasingly use discretized self-supervised speech representations for compression and integration with token-based models, yet their impact on emotional information remains unclear. We study how residual vector…

声音 · 计算机科学 2026-03-24 Haoguang Zhou , Siyi Wang , Jingyao Wu , James Bailey , Ting Dang

Any-to-any voice conversion problem aims to convert voices for source and target speakers, which are out of the training data. Previous works wildly utilize the disentangle-based models. The disentangle-based model assumes the speech…

声音 · 计算机科学 2022-02-23 Qiqi Wang , Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

The majority of existing speech emotion recognition research focuses on automatic emotion detection using training and testing data from same corpus collected under the same conditions. The performance of such systems has been shown to drop…

计算机视觉与模式识别 · 计算机科学 2020-07-29 Siddique Latif , Rajib Rana , Shahzad Younis , Junaid Qadir , Julien Epps

We propose a speech enhancement system that combines speaker-agnostic speech restoration with voice conversion (VC) to obtain a studio-level quality speech signal. While voice conversion models are typically used to change speaker…

声音 · 计算机科学 2025-05-22 Kyungguen Byun , Jason Filos , Erik Visser , Sunkuk Moon

Recently, voice conversion (VC) has been widely studied. Many VC systems use disentangle-based learning techniques to separate the speaker and the linguistic content information from a speech signal. Subsequently, they convert the voice by…

音频与语音处理 · 电气工程与系统科学 2020-11-03 Yen-Hao Chen , Da-Yi Wu , Tsung-Han Wu , Hung-yi Lee

Non-speech emotion recognition has a wide range of applications including healthcare, crime control and rescue, and entertainment, to name a few. Providing these applications using edge computing has great potential, however, recent studies…

声音 · 计算机科学 2023-05-02 Ibrahim Malik , Siddique Latif , Sanaullah Manzoor , Muhammad Usama , Junaid Qadir , Raja Jurdak

Acquiring a large vocabulary is an important aspect of human intelligence. Onecommon approach for human to populating vocabulary is to learn words duringreading or listening, and then use them in writing or speaking. This ability totransfer…

人工智能 · 计算机科学 2018-11-05 Yuanpeng Li , Yi Yang , Jianyu Wang , Wei Xu

This paper shows that StarGAN-VC, a spectral envelope transformation method for non-parallel many-to-many voice conversion (VC), is capable of emotional VC (EVC). Although StarGAN-VC has been shown to enable speaker identity conversion, its…

声音 · 计算机科学 2021-04-06 Asuka Moritani , Ryo Ozaki , Shoki Sakamoto , Hirokazu Kameoka , Tadahiro Taniguchi

In real-world singing voice conversion (SVC) applications, environmental noise and the demand for expressive output pose significant challenges. Conventional methods, however, are typically designed without accounting for real deployment…

声音 · 计算机科学 2025-10-24 Junjie Zheng , Gongyu Chen , Chaofan Ding , Zihao Chen

While most research into speech synthesis has focused on synthesizing high-quality speech for in-dataset speakers, an equally essential yet unsolved problem is synthesizing speech for unseen speakers who are out-of-dataset with limited…

声音 · 计算机科学 2023-08-28 Wenbin Wang , Yang Song , Sanjay Jha

The zero-shot scenario for speech generation aims at synthesizing a novel unseen voice with only one utterance of the target speaker. Although the challenges of adapting new voices in zero-shot scenario exist in both stages -- acoustic…

声音 · 计算机科学 2022-07-06 Yi Lei , Shan Yang , Jian Cong , Lei Xie , Dan Su

A general disentanglement-based speaker anonymization system typically separates speech into content, speaker, and prosody features using individual encoders. This paper explores how to adapt such a system when a new speech attribute, for…

Most existing Zero-Shot Text-To-Speech(ZS-TTS) systems generate the unseen speech based on single prompt, such as reference speech or text descriptions, which limits their flexibility. We propose a customized emotion ZS-TTS system based on…

声音 · 计算机科学 2025-05-27 Zhichao Wu , Yueteng Kang , Songjun Cao , Long Ma , Qiulin Li , Qun Yang

In this project, we aim to build a Text-to-Speech system able to produce speech with a controllable emotional expressiveness. We propose a methodology for solving this problem in three main steps. The first is the collection of emotional…

音频与语音处理 · 电气工程与系统科学 2019-07-08 Noé Tits

This paper introduces a new framework for non-parallel emotion conversion in speech. Our framework is based on two key contributions. First, we propose a stochastic version of the popular CycleGAN model. Our modified loss function…

音频与语音处理 · 电气工程与系统科学 2022-11-10 Ravi Shankar , Hsi-Wei Hsieh , Nicolas Charon , Archana Venkataraman

Recent developments in neural speech synthesis and vocoding have sparked a renewed interest in voice conversion (VC). Beyond timbre transfer, achieving controllability on para-linguistic parameters such as pitch and Speed is critical in…

音频与语音处理 · 电气工程与系统科学 2025-04-16 Meiying Chen , Zhiyao Duan

Real-time EEG-based Emotion Recognition (EEG-ER) with consumer-grade EEG devices involves classification of emotions using a reduced number of channels. These devices typically provide only four or five channels, unlike the high number of…

机器学习 · 计算机科学 2021-11-15 Josef Bajada , Francesco Borg Bonello

Precise control over speech characteristics, such as pitch, duration, and speech rate, remains a significant challenge in the field of voice conversion. The ability to manipulate parameters like pitch and syllable rate is an important…

声音 · 计算机科学 2025-07-08 Mathilde Abrassart , Nicolas Obin , Axel Roebel
‹ 上一页 1 8 9 10 下一页 ›