English
Related papers

Related papers: BandCondiNet: Parallel Transformers-based Conditio…

200 papers

High-level musical qualities (such as emotion) are often abstract, subjective, and hard to quantify. Given these difficulties, it is not easy to learn good feature representations with supervised learning techniques, either because of the…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-31 Hao Hao Tan , Dorien Herremans

In this paper, we propose a novel conditional convolution network, named location-variable convolution, to model the dependencies of the waveform sequence. Different from the use of unified convolution kernels in WaveNet to capture the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-02-23 Zhen Zeng , Jianzong Wang , Ning Cheng , Jing Xiao

We introduce VampNet, a masked acoustic token modeling approach to music synthesis, compression, inpainting, and variation. We use a variable masking schedule during training which allows us to sample coherent music from the model by…

Sound · Computer Science 2023-07-13 Hugo Flores Garcia , Prem Seetharaman , Rithesh Kumar , Bryan Pardo

To enhance the controllability of text-to-image diffusion models, existing efforts like ControlNet incorporated image-based conditional controls. In this paper, we reveal that existing methods still face significant challenges in generating…

Computer Vision and Pattern Recognition · Computer Science 2024-11-20 Ming Li , Taojiannan Yang , Huafeng Kuang , Jie Wu , Zhaoning Wang , Xuefeng Xiao , Chen Chen

Diffusion models have shown promising results in cross-modal generation tasks, including text-to-image and text-to-audio generation. However, generating music, as a special type of audio, presents unique challenges due to limited…

Sound · Computer Science 2023-08-04 Ke Chen , Yusong Wu , Haohe Liu , Marianna Nezhurina , Taylor Berg-Kirkpatrick , Shlomo Dubnov

Generating sound effects for product-level videos, where only a small amount of labeled data is available for diverse scenes, requires the production of high-quality sounds in few-shot settings. To tackle the challenge of limited labeled…

Generating images conditioned on multiple visual references is critical for real-world applications such as multi-subject composition, narrative illustration, and novel view synthesis, yet current models suffer from severe performance…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Zhekai Chen , Yuqing Wang , Manyuan Zhang , Xihui Liu

In this work, we provide a comprehensive survey of AI music generation tools, including both research projects and commercialized applications. To conduct our analysis, we classified music generation approaches into three categories:…

Sound · Computer Science 2023-08-28 Yueyue Zhu , Jared Baca , Banafsheh Rekabdar , Reza Rawassizadeh

Conditional generation of time-dependent data is a task that has much interest, whether for data augmentation, scenario simulation, completing missing data, or other purposes. Recent works proposed a Transformer-based Time series generative…

Machine Learning · Computer Science 2022-10-06 Abdellah Madane , Mohamed-djallel Dilmi , Florent Forest , Hanane Azzag , Mustapha Lebbah , Jerome Lacaille

With the rise of AI-generated content (AIGC), generating perceptually natural and feeling-aligned music from multimodal inputs has become a central challenge. Existing approaches often rely on explicit emotion labels that require costly…

Sound · Computer Science 2025-12-02 Jiaying Hong , Ting Zhu , Thanet Markchom , Huizhi Liang

We consider a novel task of automatically generating text descriptions of music. Compared with other well-established text generation tasks such as image caption, the scarcity of well-paired music and text datasets makes it a much more…

Sound · Computer Science 2022-09-07 Peining Zhang , Junliang Guo , Linli Xu , Mu You , Junming Yin

Conditional diffusion models have exhibited superior performance in high-fidelity text-guided visual generation and editing. Nevertheless, prevailing text-guided visual diffusion models primarily focus on incorporating text-visual…

Computer Vision and Pattern Recognition · Computer Science 2024-06-05 Ling Yang , Zhilong Zhang , Zhaochen Yu , Jingwei Liu , Minkai Xu , Stefano Ermon , Bin Cui

With the rapid development of diffusion models in image generation, the demand for more powerful and flexible controllable frameworks is increasing. Although existing methods can guide generation beyond text prompts, the challenge of…

Computer Vision and Pattern Recognition · Computer Science 2025-07-09 Haoxuan Wang , Jinlong Peng , Qingdong He , Hao Yang , Ying Jin , Jiafu Wu , Xiaobin Hu , Yanjie Pan , Zhenye Gan , Mingmin Chi , Bo Peng , Yabiao Wang

Fine-tuning large-scale music generation models, such as MusicGen and Mustango, is a computationally expensive process, often requiring updates to billions of parameters and, therefore, significant hardware resources. Parameter-Efficient…

Sound · Computer Science 2025-08-12 Atharva Mehta , Shivam Chauhan , Monojit Choudhury

Recent advances in diffusion-based text-to-image generation have demonstrated promising results through visual condition control. However, existing ControlNet-like methods struggle with compositional visual conditioning - simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Yanjie Pan , Qingdong He , Zhengkai Jiang , Pengcheng Xu , Chaoyi Wang , Jinlong Peng , Haoxuan Wang , Yun Cao , Zhenye Gan , Mingmin Chi , Bo Peng , Yabiao Wang

Automatic transcription of acoustic guitar fingerpicking performances remains a challenging task due to the scarcity of labeled training data and legal constraints connected with musical recordings. This work investigates a procedural data…

Sound · Computer Science 2025-08-12 Sebastian Murgul , Michael Heizmann

This paper makes several contributions to automatic lyrics transcription (ALT) research. Our main contribution is a novel variant of the Multistreaming Time-Delay Neural Network (MTDNN) architecture, called MSTRE-Net, which processes the…

Sound · Computer Science 2021-08-06 Emir Demirel , Sven Ahlbäck , Simon Dixon

AI creation, such as poem or lyrics generation, has attracted increasing attention from both industry and academic communities, with many promising models proposed in the past few years. Existing methods usually estimate the outputs based…

Artificial Intelligence · Computer Science 2024-09-05 Qian Cao , Xu Chen , Ruihua Song , Hao Jiang , Guang Yang , Zhao Cao

In the realm of music recommendation, sequential recommender systems have shown promise in capturing the dynamic nature of music consumption. Nevertheless, traditional Transformer-based models, such as SASRec and BERT4Rec, while effective,…

Information Retrieval · Computer Science 2024-09-09 Davide Abbattista , Vito Walter Anelli , Tommaso Di Noia , Craig Macdonald , Aleksandr Vladimirovich Petrov

Chord recognition serves as a critical task in music information retrieval due to the abstract and descriptive nature of chords in music analysis. While audio chord recognition systems have achieved significant accuracy for small…

‹ Prev 1 8 9 10 Next ›