English
Related papers

Related papers: Latent Space Explorations of Singing Voice Synthes…

200 papers

High-fidelity multi-singer singing voice synthesis is challenging for neural vocoder due to the singing voice data shortage, limited singer generalization, and large computational cost. Existing open corpora could not meet requirements for…

Audio and Speech Processing · Electrical Eng. & Systems 2021-12-21 Rongjie Huang , Feiyang Chen , Yi Ren , Jinglin Liu , Chenye Cui , Zhou Zhao

Previous approaches in singer identification have used one of monophonic vocal tracks or mixed tracks containing multiple instruments, leaving a semantic gap between these two domains of audio. In this paper, we present a system to learn a…

Sound · Computer Science 2019-06-27 Kyungyun Lee , Juhan Nam

In singing voice synthesis (SVS), generating singing voices from musical scores faces challenges due to limited data availability. This study proposes a unique strategy to address the data scarcity in SVS. We employ an existing singing…

Sound · Computer Science 2024-06-14 Jiatong Shi , Yueqian Lin , Xinyi Bai , Keyi Zhang , Yuning Wu , Yuxun Tang , Yifeng Yu , Qin Jin , Shinji Watanabe

In this paper, we address the problem of speaker verification in conditions unseen or unknown during development. A standard method for speaker verification consists of extracting speaker embeddings with a deep neural network and processing…

Sound · Computer Science 2021-08-18 Luciana Ferrer , Mitchell McLaren , Niko Brummer

Singing voice separation attempts to separate the vocal and instrumental parts of a music recording, which is a fundamental problem in music information retrieval. Recent work on singing voice separation has shown that the low-rank…

Audio and Speech Processing · Electrical Eng. & Systems 2018-01-12 Tak-Shing T. Chan , Yi-Hsuan Yang

We introduce a new version of deep state-space models (DSSMs) that combines a recurrent neural network with a state-space framework to forecast time series data. The model estimates the observed series as functions of latent variables that…

Machine Learning · Statistics 2022-05-20 Haoxuan Wu , David S. Matteson , Martin T. Wells

Zero-shot singing voice synthesis (SVS) with style transfer and style control aims to generate high-quality singing voices with unseen timbres and styles (including singing method, emotion, rhythm, technique, and pronunciation) from audio…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-02 Yu Zhang , Ziyue Jiang , Ruiqi Li , Changhao Pan , Jinzheng He , Rongjie Huang , Chuxin Wang , Zhou Zhao

FM Synthesis is a well-known algorithm used to generate complex timbre from a compact set of design primitives. Typically featuring a MIDI interface, it is usually impractical to control it from an audio source. On the other hand,…

Sound · Computer Science 2022-08-15 Franco Caspe , Andrew McPherson , Mark Sandler

Recent progress in singing voice separation has primarily focused on supervised deep learning methods. However, the scarcity of ground-truth data with clean musical sources has been a problem for long. Given a limited set of labeled data,…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-17 Zhepei Wang , Ritwik Giri , Umut Isik , Jean-Marc Valin , Arvindh Krishnaswamy

The field of Singing Voice Synthesis (SVS) has seen significant advancements in recent years due to the rapid progress of diffusion-based approaches. However, capturing vocal style, genre-specific pitch inflections, and language-dependent…

Sound · Computer Science 2025-12-01 Sandipan Dhar , Mayank Gupta , Preeti Rao

Generative models for singing voice have been mostly concerned with the task of ``singing voice synthesis,'' i.e., to produce singing voice waveforms given musical scores and text lyrics. In this work, we explore a novel yet challenging…

Sound · Computer Science 2020-07-22 Jen-Yu Liu , Yu-Hua Chen , Yin-Cheng Yeh , Yi-Hsuan Yang

Singing voice synthesis has made remarkable progress in generating natural and high-quality voices. However, existing methods rarely provide precise control over vocal techniques such as intensity, mixed voice, falsetto, bubble, and breathy…

Sound · Computer Science 2025-04-22 Wenxiang Guo , Yu Zhang , Changhao Pan , Rongjie Huang , Li Tang , Ruiqi Li , Zhiqing Hong , Yongqi Wang , Zhou Zhao

Neural TTS has shown it can generate high quality synthesized speech. In this paper, we investigate the multi-speaker latent space to improve neural TTS for adapting the system to new speakers with only several minutes of speech or…

Audio and Speech Processing · Electrical Eng. & Systems 2019-09-04 Yan Deng , Lei He , Frank Soong

Most existing neural-based text-to-speech methods rely on extensive datasets and face challenges under low-resource condition. In this paper, we introduce a novel semi-supervised text-to-speech synthesis model that learns from both paired…

Sound · Computer Science 2024-02-05 Jianzong Wang , Pengcheng Li , Xulong Zhang , Ning Cheng , Jing Xiao

The virtual world is being established in which digital humans are created indistinguishable from real humans. Producing their audio-related capabilities is crucial since voice conveys extensive personal characteristics. We aim to create a…

Sound · Computer Science 2023-05-10 Wei Xue , Yiwen Wang , Qifeng Liu , Yike Guo

Speech enhancement models should meet very low latency requirements typically smaller than 5 ms for hearing assistive devices. While various low-latency techniques have been proposed, comparing these methods in a controlled setup using DNNs…

Audio and Speech Processing · Electrical Eng. & Systems 2024-09-17 Haibin Wu , Sebastian Braun

Discrete representation has shown advantages in speech generation tasks, wherein discrete tokens are derived by discretizing hidden features from self-supervised learning (SSL) pre-trained models. However, the direct application of speech…

Sound · Computer Science 2024-06-21 Yuxun Tang , Yuning Wu , Jiatong Shi , Qin Jin

The deepfake generation of singing vocals is a concerning issue for artists in the music industry. In this work, we propose a singing voice deepfake detection (SVDD) system, which uses noise-variant encodings of open-AI's Whisper model. As…

Sound · Computer Science 2025-02-03 Falguni Sharma , Priyanka Gupta

Deep learning based singing voice synthesis (SVS) systems have been demonstrated to flexibly generate singing with better qualities, compared to conventional statistical parametric based methods. However, neural systems are generally…

Audio and Speech Processing · Electrical Eng. & Systems 2022-07-07 Shuai Guo , Jiatong Shi , Tao Qian , Shinji Watanabe , Qin Jin

Singing voice synthesis is a generative task that involves multi-dimensional control of the singing model, including lyrics, pitch, and duration, and includes the timbre of the singer and singing skills such as vibrato. In this paper, we…

Sound · Computer Science 2022-05-25 Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao
‹ Prev 1 3 4 5 6 7 10 Next ›