English
Related papers

Related papers: DuTa-VC: A Duration-aware Typical-to-atypical Voic…

200 papers

This paper proposed a novel approach for the detection and reconstruction of dysarthric speech. The encoder-decoder model factorizes speech into a low-dimensional latent space and encoding of the input text. We showed that the latent space…

Audio and Speech Processing · Electrical Eng. & Systems 2019-07-11 Daniel Korzekwa , Roberto Barra-Chicote , Bozena Kostek , Thomas Drugman , Mateusz Lajszczak

In this paper, we propose an enhanced audio-visual deep detection method. Recent methods in audio-visual deepfake detection mostly assess the synchronization between audio and visual features. Although they have shown promising results,…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Marcella Astrid , Enjie Ghorbel , Djamila Aouada

In this paper, we describe our speech generation system for the first Audio Deep Synthesis Detection Challenge (ADD 2022). Firstly, we build an any-to-many voice conversion (VC) system to convert source speech with arbitrary language…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-21 Cheng Wen , Tingwei Guo , Xingjun Tan , Rui Yan , Shuran Zhou , Chuandong Xie , Wei Zou , Xiangang Li

Text-to-speech (TTS) methods have shown promising results in voice cloning, but they require a large number of labeled text-speech pairs. Minimally-supervised speech synthesis decouples TTS by combining two types of discrete speech…

Sound · Computer Science 2023-12-19 Chunyu Qiang , Hao Li , Yixin Tian , Yi Zhao , Ying Zhang , Longbiao Wang , Jianwu Dang

Diffusion-based voice conversion (VC) techniques such as VoiceGrad have attracted interest because of their high VC performance in terms of speech quality and speaker similarity. However, a notable limitation is the slow inference caused by…

Sound · Computer Science 2024-09-05 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Yuto Kondo

This paper presents AC-VC (Almost Causal Voice Conversion), a phonetic posteriorgrams based voice conversion system that can perform any-to-many voice conversion while having only 57.5 ms future look-ahead. The complete system is composed…

Audio and Speech Processing · Electrical Eng. & Systems 2021-11-15 Damien Ronssin , Milos Cernak

Automatic speaker recognition algorithms typically characterize speech audio using short-term spectral features that encode the physiological and anatomical aspects of speech production. Such algorithms do not fully capitalize on…

Sound · Computer Science 2021-02-16 Anurag Chowdhury , Arun Ross , Prabu David

Vocal entrainment is a social adaptation mechanism in human interaction, knowledge of which can offer useful insights to an individual's cognitive-behavioral characteristics. We propose a context-aware approach for measuring vocal…

Audio and Speech Processing · Electrical Eng. & Systems 2022-11-08 Rimita Lahiri , Md Nasir , Catherine Lord , So Hyun Kim , Shrikanth Narayanan

Diffusion Models (DMs) have achieved remarkable success in realistic voice cloning (VC), while they also increase the risk of malicious misuse. Existing proactive defenses designed for traditional VC models aim to disrupt the forgery…

Sound · Computer Science 2025-12-10 Qianyue Hu , Junyan Wu , Wei Lu , Xiangyang Luo

To realize any-to-any (A2A) voice conversion (VC), most methods are to perform symmetric self-supervised reconstruction tasks (Xi to Xi), which usually results in inefficient performances due to inadequate feature decoupling, especially for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-07 Yewei Gu , Zhenyu Zhang , Xiaowei Yi , Xianfeng Zhao

This paper describes an end-to-end adversarial singing voice conversion (EA-SVC) approach. It can directly generate arbitrary singing waveform by given phonetic posteriorgram (PPG) representing content, F0 representing pitch, and speaker…

Sound · Computer Science 2020-12-04 Haohan Guo , Heng Lu , Na Hu , Chunlei Zhang , Shan Yang , Lei Xie , Dan Su , Dong Yu

Modelling prosody variation is critical for synthesizing natural and expressive speech in end-to-end text-to-speech (TTS) systems. In this paper, a cross-utterance conditional VAE (CUC-VAE) is proposed to estimate a posterior probability…

Sound · Computer Science 2022-05-10 Yang Li , Cheng Yu , Guangzhi Sun , Hua Jiang , Fanglei Sun , Weiqin Zu , Ying Wen , Yang Yang , Jun Wang

By representing speaker characteristic as a single fixed-length vector extracted solely from speech, we can train a neural multi-speaker speech synthesis model by conditioning the model on those vectors. This model can also be adapted to…

Audio and Speech Processing · Electrical Eng. & Systems 2019-10-09 Hieu-Thi Luong , Junichi Yamagishi

The Emotional Voice Conversion (EVC) aims to convert the discrete emotional state from the source emotion to the target for a given speech utterance while preserving linguistic content. In this paper, we propose regularizing emotion…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-31 Ashishkumar Gudmalwar , Ishan D. Biyani , Nirmesh Shah , Pankaj Wasnik , Rajiv Ratn Shah

Recently, attention-based transformers have become a de facto standard in many deep learning applications including natural language processing, computer vision, signal processing, etc.. In this paper, we propose a transformer-based…

Sound · Computer Science 2024-09-04 Tathagata Bandyopadhyay

This paper proposes a new approach to duration modelling for statistical parametric speech synthesis in which a recurrent statistical model is trained to output a phone transition probability at each timestep (acoustic frame). Unlike…

Computation and Language · Computer Science 2020-07-28 Srikanth Ronanki , Oliver Watts , Simon King , Gustav Eje Henter

Direct speech-to-speech translation (S2ST) translates speech from one language into another using a single model. However, due to the presence of linguistic and acoustic diversity, the target speech follows a complex multimodal…

Computation and Language · Computer Science 2023-10-12 Qingkai Fang , Yan Zhou , Yang Feng

Voice conversion is becoming increasingly popular, and a growing number of application scenarios require models with streaming inference capabilities. The recently proposed DualVC attempts to achieve this objective through streaming model…

Audio and Speech Processing · Electrical Eng. & Systems 2024-01-19 Ziqian Ning , Yuepeng Jiang , Pengcheng Zhu , Shuai Wang , Jixun Yao , Lei Xie , Mengxiao Bi

Goal: Numerous studies had successfully differentiated normal and abnormal voice samples. Nevertheless, further classification had rarely been attempted. This study proposes a novel approach, using continuous Mandarin speech instead of a…

Audio and Speech Processing · Electrical Eng. & Systems 2022-02-23 Syu-Siang Wang , Chi-Te Wang , Chih-Chung Lai , Yu Tsao , Shih-Hau Fang

Standard probabilistic linear discriminant analysis (PLDA) for speaker recognition assumes that the sample's features (usually, i-vectors) are given by a sum of three terms: a term that depends on the speaker identity, a term that models…

Machine Learning · Computer Science 2018-01-17 Luciana Ferrer