English
Related papers

Related papers: Drum-to-Vocal Percussion Sound Conversion and Its …

200 papers

We propose TES-VC (Text-driven Environment and Speaker controllable Voice Conversion), a text-driven voice conversion framework with independent control of speaker timbre and environmental acoustics. TES-VC processes simultaneous text…

Sound · Computer Science 2025-06-16 Jiawei Jin , Zhihan Yang , Yixuan Zhou , Zhiyong Wu

Text-based voice editing (TBVE) uses synthetic output from text-to-speech (TTS) systems to replace words in an original recording. Recent work has used neural models to produce edited speech that is similar to the original speech in terms…

Sound · Computer Science 2022-10-31 Jason Fong , Yun Wang , Prabhav Agrawal , Vimal Manohar , Jilong Wu , Thilo Köhler , Qing He

Emotional Voice Conversion, or emotional VC, is a technique of converting speech from one emotion state into another one, keeping the basic linguistic information and speaker identity. Previous approaches for emotional VC need parallel data…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-09 Songxiang Liu , Yuewen Cao , Helen Meng

Singing voice conversion (SVC) is one promising technique which can enrich the way of human-computer interaction by endowing a computer the ability to produce high-fidelity and expressive singing voice. In this paper, we propose DiffSVC, an…

Audio and Speech Processing · Electrical Eng. & Systems 2021-05-31 Songxiang Liu , Yuewen Cao , Dan Su , Helen Meng

Text-to-speech synthesis (TTS) is a task to convert texts into speech. Two of the factors that have been driving TTS are the advancements of probabilistic models and latent representation learning. We propose a TTS method based on latent…

Audio and Speech Processing · Electrical Eng. & Systems 2022-12-19 Yusuke Yasuda , Tomoki Toda

Speaker anonymization systems continue to improve their ability to obfuscate the original speaker characteristics in a speech signal, but often create processing artifacts and unnatural sounding voices as a tradeoff. Many of those systems…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-23 Ünal Ege Gaznepoglu , Nils Peters

Dynamic Magnetic Resonance Imaging (MRI) of the vocal tract has become an increasingly adopted imaging modality for speech motor studies. Beyond image signals, systematic data loss, noise pollution, and audio file corruption can occur due…

Sound · Computer Science 2025-12-02 Yaxuan Li , Han Jiang , Yifei Ma , Shihua Qin , Jonghye Woo , Fangxu Xing

This paper proposes visual-text to speech (vTTS), a method for synthesizing speech from visual text (i.e., text as an image). Conventional TTS converts phonemes or characters into discrete symbols and synthesizes a speech waveform from…

Unsupervised Zero-Shot Voice Conversion (VC) aims to modify the speaker characteristic of an utterance to match an unseen target speaker without relying on parallel training data. Recently, self-supervised learning of speech representation…

Sound · Computer Science 2022-02-14 Trung Dang , Dung Tran , Peter Chin , Kazuhito Koishida

Voice conversion has emerged as a pivotal technology in numerous applications ranging from assistive communication to entertainment. In this paper, we present RT-VC, a zero-shot real-time voice conversion system that delivers ultra-low…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-13 Yisi Liu , Chenyang Wang , Hanjo Kim , Raniya Khan , Gopala Anumanchipalli

One way of expressing an environmental sound is using vocal imitations, which involve the process of replicating or mimicking the rhythm and pitch of sounds by voice. We can effectively express the features of environmental sounds, such as…

Diffusion-based voice conversion (VC) techniques such as VoiceGrad have attracted interest because of their high VC performance in terms of speech quality and speaker similarity. However, a notable limitation is the slow inference caused by…

Sound · Computer Science 2024-09-05 Takuhiro Kaneko , Hirokazu Kameoka , Kou Tanaka , Yuto Kondo

Deep learning models are mostly used in an offline inference fashion. However, this strongly limits the use of these models inside audio generation setups, as most creative workflows are based on real-time digital signal processing.…

Sound · Computer Science 2022-04-15 Antoine Caillon , Philippe Esling

We propose a speech enhancement system that combines speaker-agnostic speech restoration with voice conversion (VC) to obtain a studio-level quality speech signal. While voice conversion models are typically used to change speaker…

Sound · Computer Science 2025-05-22 Kyungguen Byun , Jason Filos , Erik Visser , Sunkuk Moon

Vocoders received renewed attention as main components in statistical parametric text-to-speech (TTS) synthesis and speech transformation systems. Even though there are vocoding techniques give almost accepted synthesized speech, their high…

Sound · Computer Science 2021-06-22 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Géza Németh

Voice Conversion (VC) aims to convert the style of a source speaker, such as timbre and pitch, to the style of any target speaker while preserving the linguistic content. However, the ground truth of the converted speech does not exist in a…

Sound · Computer Science 2025-01-06 Ziqi Liang , Xulong Zhang , Chang Liu , Xiaoyang Qu , Weifeng Zhao , Jianzong Wang

The voice conversion challenge is a bi-annual scientific event held to compare and understand different voice conversion (VC) systems built on a common dataset. In 2020, we organized the third edition of the challenge and constructed and…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-31 Yi Zhao , Wen-Chin Huang , Xiaohai Tian , Junichi Yamagishi , Rohan Kumar Das , Tomi Kinnunen , Zhenhua Ling , Tomoki Toda

Voice Conversion (VC) for unseen speakers, also known as zero-shot VC, is an attractive research topic as it enables a range of applications like voice customizing, animation production, and others. Recent work in this area made progress…

Sound · Computer Science 2022-06-01 Shijun Wang , Dimche Kostadinov , Damian Borth

Custom voice is to construct a personal speech synthesis system by adapting the source speech synthesis model to the target model through the target few recordings. The solution to constructing a custom voice is to combine an adaptive…

Sound · Computer Science 2023-01-06 Xin Yuan , Yongbing Feng , Mingming Ye , Cheng Tuo , Minghang Zhang

Automatic speech recognition (ASR) systems struggle with dysarthric speech due to high inter-speaker variability and slow speaking rates. To address this, we explore dysarthric-to-healthy speech conversion for improved ASR performance. Our…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-03 Karl El Hajal , Enno Hermann , Sevada Hovsepyan , Mathew Magimai. -Doss
‹ Prev 1 8 9 10 Next ›