English
Related papers

Related papers: SongEditor: Adapting Zero-Shot Song Generation Lan…

200 papers

We propose a unified framework for Singing Voice Synthesis (SVS) and Conversion (SVC), addressing the limitations of existing approaches in cross-domain SVS/SVC, poor output musicality, and scarcity of singing data. Our framework enables…

Sound · Computer Science 2025-01-24 Shuqi Dai , Yunyun Wang , Roger B. Dannenberg , Zeyu Jin

Recent progress in multimodal models has spurred rapid advances in audio understanding, generation, and editing. However, these capabilities are typically addressed by specialized models, leaving the development of a truly unified framework…

Recent advancements in Text-to-Song generation have enabled realistic musical content production, yet existing evaluation benchmarks lack the professional granularity to capture multi-dimensional aesthetic nuances. In this paper, we propose…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-30 Dapeng Wu , Shun Lei , Wei Tan , Guangzheng Li , Yunzhe Wang , Huaicheng Zhang , Lishi Zuo , Zhiyong Wu

Rap, a prominent genre of vocal performance, remains underexplored in vocal generation. General vocal synthesis depends on precise note and duration inputs, requiring users to have related musical knowledge, which limits flexibility. In…

Audio and Speech Processing · Electrical Eng. & Systems 2024-08-29 Ziqian Ning , Shuai Wang , Yuepeng Jiang , Jixun Yao , Lei He , Shifeng Pan , Jie Ding , Lei Xie

We introduce a novel pipeline for joint audio-visual editing that enhances the coherence between edited video and its accompanying audio. Our approach first applies state-of-the-art video editing techniques to produce the target video, then…

Multimedia · Computer Science 2026-03-18 Masato Ishii , Akio Hayakawa , Takashi Shibuya , Yuki Mitsufuji

Automatic speech editing aims to modify spoken content based on textual instructions, yet traditional cascade systems suffer from complex preprocessing pipelines and a reliance on explicit external temporal alignment. Addressing these…

Sound · Computer Science 2026-01-12 Junyang Chen , Yuhang Jia , Hui Wang , Jiaming Zhou , Yaxin Han , Mengying Feng , Yong Qin

Text-editing models have recently become a prominent alternative to seq2seq models for monolingual text-generation tasks such as grammatical error correction, simplification, and style transfer. These tasks share a common trait - they…

Style transfer for out-of-domain (OOD) singing voice synthesis (SVS) focuses on generating high-quality singing voices with unseen styles (such as timbre, emotion, pronunciation, and articulation skills) derived from reference singing voice…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-02 Yu Zhang , Rongjie Huang , Ruiqi Li , JinZheng He , Yan Xia , Feiyang Chen , Xinyu Duan , Baoxing Huai , Zhou Zhao

We present SingSong, a system that generates instrumental music to accompany input vocals, potentially offering musicians and non-musicians alike an intuitive new way to create music featuring their own voice. To accomplish this, we build…

Recently, there have been significant advancements in music generation. However, existing models primarily focus on creating modern pop songs, making it challenging to produce ancient music with distinct rhythms and styles, such as ancient…

Sound · Computer Science 2026-03-02 Jiajia Li , Jiliang Hu , Ziyi Pan , Chong Chen , Zuchao Li , Ping Wang , Lefei Zhang

Existing diffusion-based video editing models have made gorgeous advances for editing attributes of a source video over time but struggle to manipulate the motion information while preserving the original protagonist's appearance and…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Shuyuan Tu , Qi Dai , Zhi-Qi Cheng , Han Hu , Xintong Han , Zuxuan Wu , Yu-Gang Jiang

Regenerating singing voices with altered lyrics while preserving melody consistency remains challenging, as existing methods either offer limited controllability or require laborious manual alignment. We propose YingMusic-Singer-Plus, a…

Audio and Speech Processing · Electrical Eng. & Systems 2026-04-10 Chunbo Hao , Junjie Zheng , Guobin Ma , Yuepeng Jiang , Huakang Chen , Wenjie Tian , Gongyu Chen , Zihao Chen , Lei Xie

The traditional songwriting process is rather complex and this is evident in the time it takes to produce lyrics that fit the genre and form comprehensive verses. Our project aims to simplify this process with deep learning techniques, thus…

Computation and Language · Computer Science 2024-09-24 Tracy Cai , Wilson Liang , Donte Townes

Speech editing and zero-shot Text-to-Speech (TTS) share a similar generative foundation conditioned on speech prompts, yet speech editing demands far stricter local acoustic consistency with surrounding unedited content. While prior work…

Sound · Computer Science 2026-05-27 Junyang Chen , Yuhang Jia , Hui Wang , Jiaming Zhou , Yongchang Gan , Yong Qin

Automatic melody-to-lyric (M2L) generation aims to create lyrics that align with a given melody. While most previous approaches generate lyrics from scratch, revision, editing plain text draft to fit it into the melody, offers a much more…

Computation and Language · Computer Science 2025-05-05 Songyan Zhao , Bingxuan Li , Yufei Tian , Nanyun Peng

Creating a pop song melody according to pre-written lyrics is a typical practice for composers. A computational model of how lyrics are set as melodies is important for automatic composition systems, but an end-to-end lyric-to-melody model…

Audio and Speech Processing · Electrical Eng. & Systems 2023-01-05 Daiyu Zhang , Ju-Chiang Wang , Katerina Kosta , Jordan B. L. Smith , Shicen Zhou

Neural codec language models achieve impressive zero-shot Text-to-Speech (TTS) by fully imitating the acoustic characteristics of a short speech prompt, including timbre, prosody, and paralinguistic information. However, such holistic…

Sound · Computer Science 2026-01-21 Hanchen Pei , Shujie Liu , Yanqing Liu , Jianwei Yu , Yuanhang Qian , Gongping Huang , Sheng Zhao , Yan Lu

Generative Large Language Models have shown impressive in-context learning abilities, performing well across various tasks with just a prompt. Previous melody-to-lyric research has been limited by scarce high-quality aligned data and…

Computation and Language · Computer Science 2024-10-03 Hong-Hsiang Liu , Yi-Wen Liu

Recent singing-voice-synthesis (SVS) methods have achieved remarkable audio quality and naturalness, yet they lack the capability to control the style attributes of the synthesized singing explicitly. We propose Prompt-Singer, the first SVS…

Sound · Computer Science 2025-01-07 Yongqi Wang , Ruofan Hu , Rongjie Huang , Zhiqing Hong , Ruiqi Li , Wenrui Liu , Fuming You , Tao Jin , Zhou Zhao

Existing pitch curve generators face two main challenges: they often neglect singer-specific expressiveness, reducing their ability to capture individual singing styles. And they are typically developed as auxiliary modules for specific…