English
Related papers

Related papers: Conan: A Chunkwise Online Network for Zero-Shot Ad…

200 papers

In this paper we propose modifications to the neural network framework, AutoVC for the task of singing technique conversion. This includes utilising a pretrained singing technique encoder which extracts technique information, upon which a…

Sound · Computer Science 2021-11-18 Brendan O'Connor , Simon Dixon , George Fazekas

Voice conversion (VC) using deep learning technologies can now generate high quality one-to-many voices and thus has been used in some practical application fields, such as entertainment and healthcare. However, voice conversion can pose…

Sound · Computer Science 2024-05-02 Qiang Huang

This work proposes a novel method to generate realistic talking head videos using audio and visual streams. We animate a source image by transferring head motion from a driving video using a dense motion field generated using learnable…

Computer Vision and Pattern Recognition · Computer Science 2022-10-07 Madhav Agarwal , Rudrabha Mukhopadhyay , Vinay Namboodiri , C V Jawahar

Many-to-many voice conversion with non-parallel training data has seen significant progress in recent years. StarGAN-based models have been interests of voice conversion. However, most of the StarGAN-based methods only focused on voice…

Sound · Computer Science 2021-04-13 Mingjie Chen , Yanpei Shi , Thomas Hain

Training deep generative models usually requires a large amount of data. To alleviate the data collection cost, the task of zero-shot GAN adaptation aims to reuse well-trained generators to synthesize images of an unseen target domain…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Seogkyu Jeon , Bei Liu , Pilhyeon Lee , Kibeom Hong , Jianlong Fu , Hyeran Byun

This paper presents a new voice conversion model capable of transforming both speaking and singing voices. It addresses key challenges in current systems, such as conveying emotions, managing pronunciation and accent changes, and…

Sound · Computer Science 2024-12-12 Sowmya Cheripally

Recent developments in generative models have shown that deep learning combined with traditional digital signal processing (DSP) techniques could successfully generate convincing violin samples [1], that source-excitation combined with…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-07 Ollie McCarthy , Zohaib Ahmed

We propose a learning-based filter that allows us to directly modify a synthetic speech waveform into a natural speech waveform. Speech-processing systems using a vocoder framework such as statistical parametric speech synthesis and voice…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-02 Kou Tanaka , Takuhiro Kaneko , Nobukatsu Hojo , Hirokazu Kameoka

Most current speech technology systems are designed to operate well even in the presence of multiple active speakers. However, most solutions assume that the number of co-current speakers is known. Unfortunately, this information might not…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-02 Midia Yousefi , John H. L. Hansen

Currently, the chunk-wise schemes are often used to make Automatic Speech Recognition (ASR) models to support streaming deployment. However, existing approaches are unable to capture the global context, lack support for parallel training,…

Sound · Computer Science 2024-02-06 Fangyuan Wang , Bo Xu , Bo Xu

We present an approach to synthesize whisper by applying a handcrafted signal processing recipe and Voice Conversion (VC) techniques to convert normally phonated speech to whispered speech. We investigate using Gaussian Mixture Models (GMM)…

Deep learning models are mostly used in an offline inference fashion. However, this strongly limits the use of these models inside audio generation setups, as most creative workflows are based on real-time digital signal processing.…

Sound · Computer Science 2022-04-15 Antoine Caillon , Philippe Esling

Any-to-any singing voice conversion (SVC) is confronted with the challenge of ``timbre leakage'' issue caused by inadequate disentanglement between the content and the speaker timbre. To address this issue, this study introduces NeuCoSVC, a…

Sound · Computer Science 2024-01-09 Binzhu Sha , Xu Li , Zhiyong Wu , Ying Shan , Helen Meng

Voice conversion is a method that allows for the transformation of speaking style while maintaining the integrity of linguistic information. There are many researchers using deep generative models for voice conversion tasks. Generative…

Sound · Computer Science 2023-08-29 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

We present a novel high-fidelity real-time neural vocoder called VocGAN. A recently developed GAN-based vocoder, MelGAN, produces speech waveforms in real-time. However, it often produces a waveform that is insufficient in quality or…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-31 Jinhyeok Yang , Junmo Lee , Youngik Kim , Hoonyoung Cho , Injung Kim

Singing voice conversion aims to transform a source singing voice into that of a target singer while preserving the original lyrics, melody, and various vocal techniques. In this paper, we propose a high-fidelity singing voice conversion…

Sound · Computer Science 2025-01-07 Yiquan Zhou , Wenyu Wang , Hongwu Ding , Jiacheng Xu , Jihua Zhu , Xin Gao , Shihao Li

One-shot style transfer is a challenging task, since training on one utterance makes model extremely easy to over-fit to training data and causes low speaker similarity and lack of expressiveness. In this paper, we build on the…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-22 Zhichao Wang , Qicong Xie , Tao Li , Hongqiang Du , Lei Xie , Pengcheng Zhu , Mengxiao Bi

Transformer has achieved competitive performance against state-of-the-art end-to-end models in automatic speech recognition (ASR), and requires significantly less training time than RNN-based models. The original Transformer, with…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-14 Wenyong Huang , Wenchao Hu , Yu Ting Yeung , Xiao Chen

Building a voice conversion system for noisy target speakers, such as users providing noisy samples or Internet found data, is a challenging task since the use of contaminated speech in model training will apparently degrade the conversion…

Sound · Computer Science 2022-07-05 Liumeng Xue , Shan Yang , Na Hu , Dan Su , Lei Xie

Audio-visual speaker extraction has attracted increasing attention, as it removes the need for pre-registered speech and leverages the visual modality as a complement to audio. Although existing methods have achieved impressive performance,…

Multimedia · Computer Science 2026-03-03 Jiadong Wang , Ke Zhang , Xinyuan Qian , Ruijie Tao , Haizhou Li , Björn Schuller