English
Related papers

Related papers: SongEditor: Adapting Zero-Shot Song Generation Lan…

200 papers

Diffusion and flow-matching models have revolutionized automatic text-to-audio generation in recent times. These models are increasingly capable of generating high quality and faithful audio outputs capturing to speech and acoustic events.…

We propose a semi-supervised singing synthesizer, which is able to learn new voices from audio data only, without any annotations such as phonetic segmentation. Our system is an encoder-decoder model with two encoders, linguistic and…

Sound · Computer Science 2020-11-06 Jordi Bonada , Merlijn Blaauw

We present a novel approach to data-to-text generation based on iterative text editing. Our approach maximizes the completeness and semantic accuracy of the output text while leveraging the abilities of recent pre-trained models for text…

Computation and Language · Computer Science 2021-01-29 Zdeněk Kasner , Ondřej Dušek

MusicGen is a music generation language model (LM) that can be conditioned on textual descriptions and melodic features. We introduce MusicGen-Chord, which extends this capability by incorporating chord progression features. This model…

Sound · Computer Science 2024-12-03 Jongmin Jung , Andreas Jansson , Dasaem Jeong

Singing Voice Synthesis (SVS) aims to generate expressive vocal performances from structured musical inputs such as lyrics and pitch sequences. While recent progress in discrete codec-based speech synthesis has enabled zero-shot generation…

Sound · Computer Science 2026-04-14 Junchuan Zhao , Wei Zeng , Tianle Lyu , Ye Wang

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee

We present Step-Audio-EditX, the first open-source LLM-based audio model excelling at expressive and iterative audio editing encompassing emotion, speaking style, and paralinguistics alongside robust zero-shot text-to-speech (TTS)…

Computation and Language · Computer Science 2025-11-20 Chao Yan , Boyong Wu , Peng Yang , Pengfei Tan , Guoqiang Hu , Li Xie , Yuxin Zhang , Xiangyu , Zhang , Fei Tian , Xuerui Yang , Xiangyu Zhang , Daxin Jiang , Shuchang Zhou , Gang Yu

Recent advancements in generative models have significantly enhanced talking face video generation, yet singing video generation remains underexplored. The differences between human talking and singing limit the performance of existing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-05 Yan Li , Ziya Zhou , Zhiqiang Wang , Wei Xue , Wenhan Luo , Yike Guo

Recent advances in singing voice synthesis (SVS) have attracted substantial attention from both academia and industry. With the advent of large language models and novel generative paradigms, producing controllable, high-fidelity singing…

Audio and Speech Processing · Electrical Eng. & Systems 2026-01-22 Changhao Pan , Dongyu Yao , Yu Zhang , Wenxiang Guo , Jingyu Lu , Zhiyuan Zhu , Zhou Zhao

Recent advancements in 3D diffusion-based semantic scene generation have gained attention. However, existing methods rely on unconditional generation and require multiple resampling steps when editing scenes, which significantly limits…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Haowen Zheng , Yanyan Liang

This paper explores the modeling method of polyphonic music sequence. Due to the great potential of Transformer models in music generation, controllable music generation is receiving more attention. In the task of polyphonic music, current…

Sound · Computer Science 2023-11-29 Jiuyang Zhou , Tengfei Niu , Hong Zhu , Xingping Wang

Pretrained language models have been shown to be effective in many software-related generation tasks; however, they are not well-suited for editing tasks as they are not designed to reason about edits. To address this, we propose a novel…

Software Engineering · Computer Science 2022-09-15 Jiyang Zhang , Sheena Panthaplackel , Pengyu Nie , Junyi Jessy Li , Milos Gligoric

Style control has been popular in video generation models. Existing methods often generate videos far from the given style, cause content leakage, and struggle to transfer one video to the desired style. Our first observation is that the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Zixuan Ye , Huijuan Huang , Xintao Wang , Pengfei Wan , Di Zhang , Wenhan Luo

We propose a sequence-to-sequence singing synthesizer, which avoids the need for training data with pre-aligned phonetic and acoustic features. Rather than the more common approach of a content-based attention mechanism combined with an…

Sound · Computer Science 2020-02-21 Merlijn Blaauw , Jordi Bonada

Creating music is iterative, requiring varied methods at each stage. However, existing AI music systems fall short in orchestrating multiple subsystems for diverse needs. To address this gap, we introduce Loop Copilot, a novel system that…

Sound · Computer Science 2024-09-02 Yixiao Zhang , Akira Maezawa , Gus Xia , Kazuhiko Yamamoto , Simon Dixon

Singing voice synthesis (SVS) is a task that aims to generate audio signals according to musical scores and lyrics. With its multifaceted nature concerning music and language, producing singing voices indistinguishable from that of human…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-07 Yin-Ping Cho , Fu-Rong Yang , Yung-Chuan Chang , Ching-Ting Cheng , Xiao-Han Wang , Yi-Wen Liu

Lifelong learning enables large language models (LLMs) to adapt to evolving information by continually updating their internal knowledge. An ideal system should support efficient, wide-ranging updates while preserving existing capabilities…

Computation and Language · Computer Science 2026-03-11 Xiaojie Gu , Ziying Huang , Jia-Chen Gu , Kai Zhang

Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited…

Sound · Computer Science 2023-08-04 Ke Chen , Yusong Wu , Haohe Liu , Marianna Nezhurina , Taylor Berg-Kirkpatrick , Shlomo Dubnov

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

Multimedia · Computer Science 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

Discrete representation has shown advantages in speech generation tasks, wherein discrete tokens are derived by discretizing hidden features from self-supervised learning (SSL) pre-trained models. However, the direct application of speech…

Sound · Computer Science 2024-06-21 Yuxun Tang , Yuning Wu , Jiatong Shi , Qin Jin