中文
相关论文

相关论文: DurFlex-EVC: Duration-Flexible Emotional Voice Con…

200 篇论文

Multimodal Large Language Models (MLLMs) excel in Open-Vocabulary (OV) emotion recognition but often neglect fine-grained acoustic modeling. Existing methods typically use global audio encoders, failing to capture subtle, local temporal…

多媒体 · 计算机科学 2026-03-24 Liyun Zhang , Xuanmeng Sha , Shuqiong Wu , Fengkai Liu

Voice conversion refers to transferring speaker identity with well-preserved content. Better disentanglement of speech representations leads to better voice conversion. Recent studies have found that phonetic information from input audio…

声音 · 计算机科学 2024-01-19 Yimin Deng , Huaizhen Tang , Xulong Zhang , Ning Cheng , Jing Xiao , Jianzong Wang

Affect is an emotional characteristic encompassing valence, arousal, and intensity, and is a crucial attribute for enabling authentic conversations. While existing text-to-speech (TTS) and speech-to-speech systems rely on strength embedding…

We introduce DISSC, a novel, lightweight method that converts the rhythm, pitch contour and timbre of a recording to a target speaker in a textless manner. Unlike DISSC, most voice conversion (VC) methods focus primarily on timbre, and…

声音 · 计算机科学 2023-10-20 Gallil Maimon , Yossi Adi

Emotion Recognition in Conversations (ERC) is an important and active research area. Recent work has shown the benefits of using multiple modalities (e.g., text, audio, and video) for the ERC task. In a conversation, participants tend to…

计算与语言 · 计算机科学 2022-11-08 Harsh Agarwal , Keshav Bansal , Abhinav Joshi , Ashutosh Modi

Accurate speech emotion recognition is essential for developing human-facing systems. Recent advancements have included finetuning large, pretrained transformer models like Wav2Vec 2.0. However, the finetuning process requires substantial…

声音 · 计算机科学 2025-03-07 Aneesha Sampath , James Tavernor , Emily Mower Provost

Data augmentation via voice conversion (VC) has been successfully applied to low-resource expressive text-to-speech (TTS) when only neutral data for the target speaker are available. Although the quality of VC is crucial for this approach,…

音频与语音处理 · 电气工程与系统科学 2022-07-06 Ryo Terashima , Ryuichi Yamamoto , Eunwoo Song , Yuma Shirahata , Hyun-Wook Yoon , Jae-Min Kim , Kentaro Tachibana

We propose SelfVC, a training strategy to iteratively improve a voice conversion model with self-synthesized examples. Previous efforts on voice conversion focus on factorizing speech into explicitly disentangled representations that…

Speech Emotion Conversion aims to modify the emotion expressed in input speech while preserving lexical content and speaker identity. Recently, generative modeling approaches have shown promising results in changing local acoustic…

音频与语音处理 · 电气工程与系统科学 2025-08-18 Navin Raj Prabhu , Danilo de Oliveira , Nale Lehmann-Willenbrock , Timo Gerkmann

Emotion Recognition in Conversations (ERC) is essential for building empathetic human-machine systems. Existing studies on ERC primarily focus on summarizing the context information in a conversation, however, ignoring the differentiated…

计算与语言 · 计算机科学 2020-10-16 Yuzhao Mao , Qi Sun , Guang Liu , Xiaojie Wang , Weiguo Gao , Xuan Li , Jianping Shen

Explicit duration modeling is a key to achieving robust and efficient alignment in text-to-speech synthesis (TTS). We propose a new TTS framework using explicit duration modeling that incorporates duration as a discrete latent variable to…

音频与语音处理 · 电气工程与系统科学 2020-10-21 Yusuke Yasuda , Xin Wang , Junichi Yamagishi

Despite the recent advances in applying pre-trained language models to generate high-quality texts, generating long passages that maintain long-range coherence is yet challenging for these models. In this paper, we propose DiscoDVT, a…

计算与语言 · 计算机科学 2021-10-13 Haozhe Ji , Minlie Huang

Recent advances in emotional voice conversion (EVC) have enabled the generation of expressive synthetic speech, raising new concerns in audio deepfake detection. Existing approaches treat speech as a homogeneous signal and largely overlook…

声音 · 计算机科学 2026-05-06 Vamshi Nallaguntla , Shruti Kshirsagar , Anderson R. Avila

Currently, zero-shot voice conversion systems are capable of synthesizing the voice of unseen speakers. However, most existing approaches struggle to accurately replicate the speaking style of the source speaker or mimic the distinctive…

声音 · 计算机科学 2025-06-02 Kaidi Wang , Wenhao Guan , Ziyue Jiang , Hukai Huang , Peijie Chen , Weijie Wu , Qingyang Hong , Lin Li

Whispered speech lacks vocal-fold excitation, making intelligible conversion challenging. We propose WhisperVC, a three-stage framework for low-resource whisper-to-normal (W2N) conversion that decouples cross-domain alignment from speech…

音频与语音处理 · 电气工程与系统科学 2026-03-11 Dong Liu , Juan Liu , Wei Ju , Yao Tian , Ming Li

We introduce SEDTalker, an emotion-aware framework for speech-driven 3D facial animation that leverages frame-level speech emotion diarization to achieve fine-grained expressive control. Unlike prior approaches that rely on utterance-level…

计算机视觉与模式识别 · 计算机科学 2026-04-16 Farzaneh Jafari , Stefano Berretti , Anup Basu

Despite advances in deep probabilistic models, learning discrete latent representations remains challenging. This work introduces a novel method to improve inference in discrete Variational Autoencoders by reframing the inference problem…

机器学习 · 计算机科学 2025-06-11 María Martínez-García , Grace Villacrés , David Mitchell , Pablo M. Olmos

In this paper, we propose a new differentiable neural network alignment mechanism for text-dependent speaker verification which uses alignment models to produce a supervector representation of an utterance. Unlike previous works with…

声音 · 计算机科学 2018-12-27 Victoria Mingote , Antonio Miguel , Alfonso Ortega , Eduardo Lleida

There are individual differences in expressive behaviors driven by cultural norms and personality. This between-person variation can result in reduced emotion recognition performance. Therefore, personalization is an important step in…

音频与语音处理 · 电气工程与系统科学 2023-09-06 Minh Tran , Yufeng Yin , Mohammad Soleymani

Text-based speech editing allows users to edit speech by intuitively cutting, copying, and pasting text to speed up the process of editing speech. In the previous work, CampNet (context-aware mask prediction network) is proposed to realize…

声音 · 计算机科学 2022-12-21 Tao Wang , Jiangyan Yi , Ruibo Fu , Jianhua Tao , Zhengqi Wen , Chu Yuan Zhang