English
Related papers

Related papers: SongCreator: Lyrics-based Universal Song Generatio…

200 papers

Large Language Models (LLMs) show promise in lyric-to-melody generation, but models trained with Supervised Fine-Tuning (SFT) often produce musically implausible melodies with issues like poor rhythm and unsuitable vocal ranges, a…

Sound · Computer Science 2026-04-21 Hao Meng , Siyuan Zheng , Shuran Zhou , Qiangqiang Wang , Yang Song

Recent advances in deep learning have expanded possibilities to generate music, but generating a customizable full piece of music with consistent long-term structure remains a challenge. This paper introduces MusicFrameworks, a hierarchical…

Sound · Computer Science 2021-09-03 Shuqi Dai , Zeyu Jin , Celso Gomes , Roger B. Dannenberg

In this paper, we develop DeepSinger, a multi-lingual multi-singer singing voice synthesis (SVS) system, which is built from scratch using singing training data mined from music websites. The pipeline of DeepSinger consists of several…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-16 Yi Ren , Xu Tan , Tao Qin , Jian Luan , Zhou Zhao , Tie-Yan Liu

Large Language Models (LLM) have shown encouraging progress in multimodal understanding and generation tasks. However, how to design a human-aligned and interpretable melody composition system is still under-explored. To solve this problem,…

Sound · Computer Science 2024-03-08 Xia Liang , Xingjian Du , Jiaju Lin , Pei Zou , Yuan Wan , Bilei Zhu

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-30 Dapeng Wu , Shun Lei , Wei Tan , Guangzheng Li , Yunzhe Wang , Huaicheng Zhang , Lishi Zuo , Zhiyong Wu

Music structure analysis (MSA) underpins music understanding and controllable generation, yet progress has been limited by small, inconsistent corpora. We present SongFormer, a scalable framework that learns from heterogeneous supervision.…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-09 Chunbo Hao , Ruibin Yuan , Jixun Yao , Qixin Deng , Xinyi Bai , Yanbo Wang , Wei Xue , Lei Xie

Singing voice synthesis (SVS) and singing voice conversion (SVC) have achieved remarkable progress in generating natural-sounding human singing. However, existing systems are restricted to human timbres and have limited ability to…

Sound · Computer Science 2025-11-27 Jionghao Han , Jiatong Shi , Zhuoyan Tao , Yuxun Tang , Yiwen Zhao , Gus Xia , Shinji Watanabe

Deep learning-based works for singing voice separation have performed exceptionally well in the recent past. However, most of these works do not focus on allowing users to interact with the model to improve performance. This can be crucial…

Sound · Computer Science 2025-12-03 Ankur Gupta , Anshul Rai , Archit Bansal , Vipul Arora

Singing, as a common facial movement second only to talking, can be regarded as a universal language across ethnicities and cultures, plays an important role in emotional communication, art, and entertainment. However, it is often…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Sijing Wu , Yunhao Li , Weitian Zhang , Jun Jia , Yucheng Zhu , Yichao Yan , Guangtao Zhai , Xiaokang Yang

Language models pretrained on text-only corpora often struggle with tasks that require auditory commonsense knowledge. Previous work addresses this problem by augmenting the language model to retrieve knowledge from external audio…

Computation and Language · Computer Science 2025-06-10 Suho Yoo , Hyunjong Ok , Jaeho Lee

Large Language Models (LLMs) have recently garnered significant attention, primarily for their capabilities in text-based interactions. However, natural human interaction often relies on speech, necessitating a shift towards voice-based…

Computation and Language · Computer Science 2025-08-08 Wenqian Cui , Dianzhi Yu , Xiaoqi Jiao , Ziqiao Meng , Guangyan Zhang , Qichao Wang , Yiwen Guo , Irwin King

We introduce AudioLM, a framework for high-quality audio generation with long-term consistency. AudioLM maps the input audio to a sequence of discrete tokens and casts audio generation as a language modeling task in this representation…

Recent deep music generation studies have put much emphasis on long-term generation with structures. However, we are yet to see high-quality, well-structured whole-song generation. In this paper, we make the first attempt to model a full…

Sound · Computer Science 2024-05-17 Ziyu Wang , Lejun Min , Gus Xia

Singing-driven 3D head animation is a challenging yet promising task with applications in virtual avatars, entertainment, and education. Unlike speech, singing involves richer emotional nuance, dynamic prosody, and lyric-based semantics,…

Graphics · Computer Science 2025-09-03 Zikai Huang , Yihan Zhou , Xuemiao Xu , Cheng Xu , Xiaofen Xing , Jing Qin , Shengfeng He

With the rapid development of neural network architectures and speech processing models, singing voice synthesis with neural networks is becoming the cutting-edge technique of digital music production. In this work, in order to explore how…

Sound · Computer Science 2021-08-29 Dengfeng Ke , Yuxing Lu , Xudong Liu , Yanyan Xu , Jing Sun , Cheng-Hao Cai

Music has a unique and complex structure which is challenging for both expert humans and existing AI systems to understand, and presents unique challenges relative to other forms of audio. We present LLark, an instruction-tuned multimodal…

Sound · Computer Science 2024-06-04 Josh Gardner , Simon Durand , Daniel Stoller , Rachel M. Bittner

The current landscape of research leveraging large language models (LLMs) is experiencing a surge. Many works harness the powerful reasoning capabilities of these models to comprehend various modalities, such as text, speech, images,…

Sound · Computer Science 2024-12-10 Shansong Liu , Atin Sakkeer Hussain , Qilong Wu , Chenshuo Sun , Ying Shan

Creating a pop song melody according to pre-written lyrics is a typical practice for composers. A computational model of how lyrics are set as melodies is important for automatic composition systems, but an end-to-end lyric-to-melody model…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-05 Daiyu Zhang , Ju-Chiang Wang , Katerina Kosta , Jordan B. L. Smith , Shicen Zhou

Large Language Models (LLMs) have revolutionised the field of Natural Language Processing (NLP) and have achieved state-of-the-art performance in practically every task in this field. However, the prevalent approach used in text generation,…

Computation and Language · Computer Science 2024-08-12 Nicolo Micheletti , Samuel Belkadi , Lifeng Han , Goran Nenadic

Audio is an essential part of our life, but creating it often requires expertise and is time-consuming. Research communities have made great progress over the past year advancing the performance of large scale audio generative models for a…