English
Related papers

Related papers: StarGAN-based Emotional Voice Conversion for Japan…

200 papers

In Emotion Recognition in Conversations (ERC), the emotions of target utterances are closely dependent on their context. Therefore, existing works train the model to generate the response of the target utterance, which aims to recognise…

Computation and Language · Computer Science 2023-05-30 Kailai Yang , Tianlin Zhang , Sophia Ananiadou

To establish empathy with machines, it is essential to fully understand human emotional changes. However, research in multimodal emotion recognition often overlooks one problem: individual expressive traits vary significantly, which means…

Sound · Computer Science 2026-04-29 Kexue Wang , Yinfeng Yu , Liejun Wang

In this work, we propose a zero-shot voice conversion method using speech representations trained with self-supervised learning. First, we develop a multi-task model to decompose a speech utterance into features such as linguistic content,…

Sound · Computer Science 2023-02-17 Shehzeen Hussain , Paarth Neekhara , Jocelyn Huang , Jason Li , Boris Ginsburg

Facial expression recognition is a challenging task due to two major problems: the presence of inter-subject variations in facial expression recognition dataset and impure expressions posed by human subjects. In this paper we present a…

Computer Vision and Pattern Recognition · Computer Science 2019-10-15 Kamran Ali , Ilkin Isler , Charles Hughes

We introduce LinearVC, a simple voice conversion method that sheds light on the structure of self-supervised representations. First, we show that simple linear transformations of self-supervised features effectively convert voices. Next, we…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Herman Kamper , Benjamin van Niekerk , Julian Zaïdi , Marc-André Carbonneau

Voice conversion (VC) systems are widely used for several applications, from speaker anonymisation to personalised speech synthesis. Supervised approaches learn a mapping between different speakers using parallel data, which is expensive to…

Despite remarkable success in image-to-image translation that celebrates the advancements of generative adversarial networks (GANs), very limited attempts are known for video domain translation. We study the task of video-to-video…

Computer Vision and Pattern Recognition · Computer Science 2019-05-30 Michail C. Doukas , Viktoriia Sharmanska , Stefanos Zafeiriou

Foreign accent conversion (FAC) is a special application of voice conversion (VC) which aims to convert the accented speech of a non-native speaker to a native-sounding speech with the same speaker identity. FAC is difficult since the…

Sound · Computer Science 2023-09-06 Wen-Chin Huang , Tomoki Toda

Emotional voice conversion models adapt the emotion in speech without changing the speaker identity or linguistic content. They are less data hungry than text-to-speech models and allow to generate large amounts of emotional data for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-01-15 Bastian Schnell , Goeric Huybrechts , Bartek Perz , Thomas Drugman , Jaime Lorenzo-Trueba

This paper proposes ESTVocoder, a novel excitation-spectral-transformed neural vocoder within the framework of source-filter theory. The ESTVocoder transforms the amplitude and phase spectra of the excitation into the corresponding speech…

Sound · Computer Science 2024-11-19 Xiao-Hang Jiang , Hui-Peng Du , Yang Ai , Ye-Xin Lu , Zhen-Hua Ling

We introduce a novel method for emotion conversion in speech that does not require parallel training data. Our approach loosely relies on a cycle-GAN schema to minimize the reconstruction error from converting back and forth between emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-12 Ravi Shankar , Jacob Sager , Archana Venkataraman

Despite the significant advancements in Text-to-Speech (TTS) systems, their full utilization in automatic dubbing remains limited. This task necessitates the extraction of voice identity and emotional style from a reference speech in a…

Audio and Speech Processing · Electrical Eng. & Systems 2024-06-13 Ashishkumar Gudmalwar , Nirmesh Shah , Sai Akarsh , Pankaj Wasnik , Rajiv Ratn Shah

Cycle consistent generative adversarial network (CycleGAN) and variational autoencoder (VAE) based models have gained popularity in non-parallel voice conversion recently. However, they often suffer from difficult training process and…

Sound · Computer Science 2021-04-05 Tingle Li , Yichen Liu , Chenxu Hu , Hang Zhao

This paper presents a novel task, zero-shot voice conversion based on face images (zero-shot FaceVC), which aims at converting the voice characteristics of an utterance from any source speaker to a newly coming target speaker, solely…

Sound · Computer Science 2023-09-19 Zheng-Yan Sheng , Yang Ai , Yan-Nian Chen , Zhen-Hua Ling

We present the Voice Conversion Challenge 2018, designed as a follow up to the 2016 edition with the aim of providing a common framework for evaluating and comparing different state-of-the-art voice conversion (VC) systems. The objective of…

Audio and Speech Processing · Electrical Eng. & Systems 2018-04-13 Jaime Lorenzo-Trueba , Junichi Yamagishi , Tomoki Toda , Daisuke Saito , Fernando Villavicencio , Tomi Kinnunen , Zhenhua Ling

Any-to-any voice conversion aims to transform source speech into a target voice with just a few examples of the target speaker as a reference. Recent methods produce convincing conversions, but at the cost of increased complexity -- making…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-31 Matthew Baas , Benjamin van Niekerk , Herman Kamper

Emotion Recognition in Conversation (ERC) is a more challenging task than conventional text emotion recognition. It can be regarded as a personalized and interactive emotion recognition task, which is supposed to consider not only the…

Computation and Language · Computer Science 2021-01-01 Jiangnan Li , Zheng Lin , Peng Fu , Qingyi Si , Weiping Wang

Non-parallel many-to-many voice conversion remains an interesting but challenging speech processing task. Many style-transfer-inspired methods such as generative adversarial networks (GANs) and variational autoencoders (VAEs) have been…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-17 Kaizhi Qian , Zeyu Jin , Mark Hasegawa-Johnson , Gautham J. Mysore

This report introduces \texttt{EEVE-Korean-v1.0}, a Korean adaptation of large language models that exhibit remarkable capabilities across English and Korean text understanding. Building on recent highly capable but English-centric LLMs,…

Computation and Language · Computer Science 2024-02-23 Seungduk Kim , Seungtaek Choi , Myeongho Jeong

This paper describes an end-to-end adversarial singing voice conversion (EA-SVC) approach. It can directly generate arbitrary singing waveform by given phonetic posteriorgram (PPG) representing content, F0 representing pitch, and speaker…

Sound · Computer Science 2020-12-04 Haohan Guo , Heng Lu , Na Hu , Chunlei Zhang , Shan Yang , Lei Xie , Dan Su , Dong Yu