English
Related papers

Related papers: An Empirical Study on End-to-End Singing Voice Syn…

200 papers

The single-speaker singing voice synthesis (SVS) usually underperforms at pitch values that are out of the singer's vocal range or associated with limited training samples. Based on our previous work, this work proposes a…

Sound · Computer Science 2023-09-04 Shaohuan Zhou , Xu Li , Zhiyong Wu , Ying Shan , Helen Meng

Separating a song into vocal and accompaniment components is an active research topic, and recent years witnessed an increased performance from supervised training using deep learning techniques. We propose to apply the visual information…

Sound · Computer Science 2021-07-02 Bochen Li , Yuxuan Wang , Zhiyao Duan

Neural models have become ubiquitous in automatic speech recognition systems. While neural networks are typically used as acoustic models in more complex systems, recent studies have explored end-to-end speech recognition systems based on…

Computation and Language · Computer Science 2017-09-15 Yonatan Belinkov , James Glass

Various parametric representations have been proposed to model the speech signal. While the performance of such vocoders is well-known in the context of speech processing, their extrapolation to singing voice synthesis might not be…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-09 Onur Babacan , Thomas Drugman , Tuomo Raitio , Daniel Erro , Thierry Dutoit

The ideal goal of voice conversion is to convert the source speaker's speech to sound naturally like the target speaker while maintaining the linguistic content and the prosody of the source speech. However, current approaches are…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-16 Qicong Xie , Shan Yang , Yi Lei , Lei Xie , Dan Su

End-to-end neural diarization (EEND) is nowadays one of the most prominent research topics in speaker diarization. EEND presents an attractive alternative to standard cascaded diarization systems since a single system is trained at once to…

Audio and Speech Processing · Electrical Eng. & Systems 2022-06-28 Federico Landini , Alicia Lozano-Diez , Mireia Diez , Lukáš Burget

Singing voice conversion (SVC) aims to convert the voice of one singer to that of other singers while keeping the singing content and melody. On top of recent voice conversion works, we propose a novel model to steadily convert songs while…

Sound · Computer Science 2020-10-29 Zhonghao Li , Benlai Tang , Xiang Yin , Yuan Wan , Ling Xu , Chen Shen , Zejun Ma

Generative models in vision have seen rapid progress due to algorithmic improvements and the availability of high-quality image datasets. In this paper, we offer contributions in both these areas to enable similar progress in audio…

Machine Learning · Computer Science 2017-04-06 Jesse Engel , Cinjon Resnick , Adam Roberts , Sander Dieleman , Douglas Eck , Karen Simonyan , Mohammad Norouzi

To date, various speech technology systems have adopted the vocoder approach, a method for synthesizing speech waveform that shows a major role in the performance of statistical parametric speech synthesis. WaveNet one of the best models…

Sound · Computer Science 2021-06-15 Mohammed Salah Al-Radhi , Tamás Gábor Csapó , Csaba Zainkó , Géza Németh

Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-30 Ke Gu , Zhicong Wu , Peng Bai , Sitong Qiao , Zhiqi Jiang , Junchen Lu , Xiaodong Shi , Xinyuan Qian

This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal…

Audio and Speech Processing · Electrical Eng. & Systems 2023-03-16 Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda

Singing voice synthesis (SVS) aims to produce high-fidelity singing audio from music scores, requiring a detailed understanding of notes, pitch, and duration, unlike text-to-speech tasks. Although diffusion models have shown exceptional…

Sound · Computer Science 2024-10-30 Kehan Sui , Jinxu Xiang , Fang Jin

Voice cloning is a highly desired feature for personalized speech interfaces. Neural network based speech synthesis has been shown to generate high quality speech for a large number of speakers. In this paper, we introduce a neural voice…

Computation and Language · Computer Science 2018-10-15 Sercan O. Arik , Jitong Chen , Kainan Peng , Wei Ping , Yanqi Zhou

A method for musical audio synthesis using autoencoding neural networks is proposed. The autoencoder is trained to compress and reconstruct magnitude short-time Fourier transform frames. The autoencoder produces a spectrogram by activating…

Audio and Speech Processing · Electrical Eng. & Systems 2020-04-29 Joseph Colonel , Christopher Curro , Sam Keene

Generative models for singing voice have been mostly concerned with the task of ``singing voice synthesis,'' i.e., to produce singing voice waveforms given musical scores and text lyrics. In this work, we explore a novel yet challenging…

Sound · Computer Science 2020-07-22 Jen-Yu Liu , Yu-Hua Chen , Yin-Cheng Yeh , Yi-Hsuan Yang

In this paper, we study a novel task that learns to compose music from natural language. Given the lyrics as input, we propose a melody composition model that generates lyrics-conditional melody as well as the exact alignment between the…

Computation and Language · Computer Science 2018-09-13 Hangbo Bao , Shaohan Huang , Furu Wei , Lei Cui , Yu Wu , Chuanqi Tan , Songhao Piao , Ming Zhou

End-to-end neural TTS has achieved superior performance on reading style speech synthesis. However, it's still a challenge to build a high-quality conversational TTS due to the limitations of the corpus and modeling capability. This study…

Sound · Computer Science 2020-11-17 Haohan Guo , Shaofei Zhang , Frank K. Soong , Lei He , Lei Xie

Controllable Singing Voice Synthesis (SVS) aims to generate expressive singing voices reflecting user intent. While recent SVS systems achieve high audio quality, most rely on probabilistic modeling, limiting precise control over attributes…

Sound · Computer Science 2025-09-10 Yerin Ryu , Inseop Shin , Chanwoo Kim

This paper introduces a multi-scale speech style modeling method for end-to-end expressive speech synthesis. The proposed method employs a multi-scale reference encoder to extract both the global-scale utterance-level and the local-scale…

Sound · Computer Science 2021-04-09 Xiang Li , Changhe Song , Jingbei Li , Zhiyong Wu , Jia Jia , Helen Meng

Both acoustic and visual information influence human perception of speech. For this reason, the lack of audio in a video sequence determines an extremely low speech intelligibility for untrained lip readers. In this paper, we present a way…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-18 Daniel Michelsanti , Olga Slizovskaia , Gloria Haro , Emilia Gómez , Zheng-Hua Tan , Jesper Jensen