中文
相关论文

相关论文: Singing Voice Synthesis Using Deep Autoregressive …

200 篇论文

This paper presents ByteSing, a Chinese singing voice synthesis (SVS) system based on duration allocated Tacotron-like acoustic models and WaveRNN neural vocoders. Different from the conventional SVS models, the proposed ByteSing employs…

音频与语音处理 · 电气工程与系统科学 2021-01-26 Yu Gu , Xiang Yin , Yonghui Rao , Yuan Wan , Benlai Tang , Yang Zhang , Jitong Chen , Yuxuan Wang , Zejun Ma

Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing…

音频与语音处理 · 电气工程与系统科学 2026-01-22 Changhao Pan , Dongyu Yao , Yu Zhang , Wenxiang Guo , Jingyu Lu , Zhiyuan Zhu , Zhou Zhao

Artificial reverberation (AR) models play a central role in various audio applications. Therefore, estimating the AR model parameters (ARPs) of a reference reverberation is a crucial task. Although a few recent deep-learning-based…

声音 · 计算机科学 2022-07-21 Sungho Lee , Hyeong-Seok Choi , Kyogu Lee

Recent progress in deep generative models has improved the quality of neural vocoders in speech domain. However, generating a high-quality singing voice remains challenging due to a wider variety of musical expressions in pitch, loudness,…

声音 · 计算机科学 2022-10-19 Naoya Takahashi , Mayank Kumar , Singh , Yuki Mitsufuji

Singing voice synthesis (SVS) aims to generate expressive and high-quality vocals from musical scores, requiring precise modeling of pitch, duration, and articulation. While diffusion-based models have achieved remarkable success in image…

声音 · 计算机科学 2025-06-27 Kehan Sui , Jinxu Xiang , Fang Jin

Singing Accompaniment Generation (SAG), which generates instrumental music to accompany input vocals, is crucial to developing human-AI symbiotic art creation systems. The state-of-the-art method, SingSong, utilizes a multi-stage…

声音 · 计算机科学 2024-05-14 Jianyi Chen , Wei Xue , Xu Tan , Zhen Ye , Qifeng Liu , Yike Guo

This paper aims to introduce a robust singing voice synthesis (SVS) system to produce very natural and realistic singing voices efficiently by leveraging the adversarial training strategy. On one hand, we designed simple but generic random…

声音 · 计算机科学 2023-02-17 Zewang Zhang , Yibin Zheng , Xinhui Li , Li Lu

Singing voice synthesis has been paid rising attention with the rapid development of speech synthesis area. In general, a studio-level singing corpus is usually necessary to produce a natural singing voice from lyrics and music-related…

声音 · 计算机科学 2020-11-18 Heyang Xue , Shan Yang , Yi Lei , Lei Xie , Xiulin Li

Deep generative models have achieved significant progress in speech synthesis to date, while high-fidelity singing voice synthesis is still an open problem for its long continuous pronunciation, rich high-frequency parts, and strong…

音频与语音处理 · 电气工程与系统科学 2022-08-08 Rongjie Huang , Chenye Cui , Feiyang Chen , Yi Ren , Jinglin Liu , Zhou Zhao , Baoxing Huai , Zhefeng Wang

In this paper, we propose VISinger, a complete end-to-end high-quality singing voice synthesis (SVS) system that directly generates audio waveform from lyrics and musical score. Our approach is inspired by VITS, which adopts VAE-based…

音频与语音处理 · 电气工程与系统科学 2022-02-25 Yongmao Zhang , Jian Cong , Heyang Xue , Lei Xie , Pengcheng Zhu , Mengxiao Bi

We propose a sequence-to-sequence singing synthesizer, which avoids the need for training data with pre-aligned phonetic and acoustic features. Rather than the more common approach of a content-based attention mechanism combined with an…

声音 · 计算机科学 2020-02-21 Merlijn Blaauw , Jordi Bonada

This paper proposes a novel sequence-to-sequence (seq2seq) model with a musical note position-aware attention mechanism for singing voice synthesis (SVS). A seq2seq modeling approach that can simultaneously perform acoustic and temporal…

音频与语音处理 · 电气工程与系统科学 2023-03-16 Yukiya Hono , Kei Hashimoto , Yoshihiko Nankaku , Keiichi Tokuda

Recently, deep learning-based generative models have been introduced to generate singing voices. One approach is to predict the parametric vocoder features consisting of explicit speech parameters. This approach has the advantage that the…

音频与语音处理 · 电气工程与系统科学 2024-06-14 Tae-Woo Kim , Min-Su Kang , Gyeong-Hoon Lee

Existing singing voice synthesis (SVS) models largely rely on fine-grained, phoneme-level durations, which limits their practical application. These methods overlook the complementary role of visual information in duration prediction.To…

音频与语音处理 · 电气工程与系统科学 2025-09-30 Ke Gu , Zhicong Wu , Peng Bai , Sitong Qiao , Zhiqi Jiang , Junchen Lu , Xiaodong Shi , Xinyuan Qian

We investigate the feasibility of a singing voice synthesis (SVS) system by using a decomposed framework to improve flexibility in generating singing voices. Due to data-driven approaches, SVS performs a music score-to-waveform mapping;…

声音 · 计算机科学 2024-07-15 Lester Phillip Violeta , Taketo Akama

In this paper, we develop a new multi-singer Chinese neural singing voice synthesis (SVS) system named WeSinger. To improve the accuracy and naturalness of synthesized singing voice, we design several specifical modules and techniques: 1) A…

声音 · 计算机科学 2022-06-28 Zewang Zhang , Yibin Zheng , Xinhui Li , Li Lu

The objective of deep learning methods based on encoder-decoder architectures for music source separation is to approximate either ideal time-frequency masks or spectral representations of the target music source(s). The spectral…

In this paper, we propose a model to perform style transfer of speech to singing voice. Contrary to the previous signal processing-based methods, which require high-quality singing templates or phoneme synchronization, we explore a…

声音 · 计算机科学 2022-08-29 Shrutina Agarwal , Sriram Ganapathy , Naoya Takahashi

Deep learning based singing voice synthesis (SVS) systems have been demonstrated to flexibly generate singing with better qualities, compared to conventional statistical parametric based methods. However, neural systems are generally…

音频与语音处理 · 电气工程与系统科学 2022-07-07 Shuai Guo , Jiatong Shi , Tao Qian , Shinji Watanabe , Qin Jin

This paper presents a new method of singing voice analysis that performs mutually-dependent singing voice separation and vocal fundamental frequency (F0) estimation. Vocal F0 estimation is considered to become easier if singing voices can…

声音 · 计算机科学 2016-11-29 Yukara Ikemiya , Katsutoshi Itoyama , Kazuyoshi Yoshii