English
Related papers

Related papers: ViTex: Visual Texture Control for Multi-Track Symb…

200 papers

We introduce MIDI-VAE, a neural network model based on Variational Autoencoders that is capable of handling polyphonic music with multiple instrument tracks, as well as modeling the dynamics of music by incorporating note durations and…

Sound · Computer Science 2018-09-21 Gino Brunner , Andres Konrad , Yuyi Wang , Roger Wattenhofer

We demonstrate that pre-trained text-to-image diffusion models, despite being trained on raster images, possess a remarkable capacity to guide vector sketch synthesis. In this paper, we introduce DiffSketcher, a novel algorithm for…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Ximing Xing , Chuang Wang , Haitao Zhou , Jing Zhang , Qian Yu , Dong Xu

Text-to-audio (TTA) generation with fine-grained control signals, e.g., precise timing control or intelligible speech content, has been explored in recent works. However, constrained by data scarcity, their generation performance at scale…

Sound · Computer Science 2026-04-21 Yuxuan Jiang , Zehua Chen , Zeqian Ju , Yusheng Dai , Weibei Dou , Jun Zhu

With the advancement of generative artificial intelligence, previous studies have achieved the task of generating aesthetic images from hand-drawn sketches, fulfilling the public's needs for drawing. However, these methods are limited to…

Computer Vision and Pattern Recognition · Computer Science 2025-02-18 Lifan Jiang , Shuang Chen , Boxi Wu , Xiaotong Guan , Jiahui Zhang

Understanding how large audio models represent music, and using that understanding to steer generation, is both challenging and underexplored. Inspired by mechanistic interpretability in language models, where direction vectors in…

We propose MuseControlLite, a lightweight mechanism designed to fine-tune text-to-music generation models for precise conditioning using various time-varying musical attributes and reference audio signals. The key finding is that positional…

We introduce a method for composing object-level visual prompts within a text-to-image diffusion model. Our approach addresses the task of generating semantically coherent compositions across diverse scenes and styles, similar to the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Gaurav Parmar , Or Patashnik , Kuan-Chieh Wang , Daniil Ostashev , Srinivasa Narasimhan , Jun-Yan Zhu , Daniel Cohen-Or , Kfir Aberman

We present JASCO, a temporally controlled text-to-music generation model utilizing both symbolic and audio-based conditions. JASCO can generate high-quality music samples conditioned on global text descriptions along with fine-grained local…

Sound · Computer Science 2024-06-18 Or Tal , Alon Ziv , Itai Gat , Felix Kreuk , Yossi Adi

Textual image generation spans diverse fields like advertising, education, product packaging, social media, information visualization, and branding. Despite recent strides in language-guided image synthesis using diffusion models, current…

Computer Vision and Pattern Recognition · Computer Science 2024-05-22 Shubham Paliwal , Arushi Jain , Monika Sharma , Vikram Jamwal , Lovekesh Vig

Recently, multi-instrument music generation has become a hot topic. Different from single-instrument generation, multi-instrument generation needs to consider inter-track harmony besides intra-track coherence. This is usually achieved by…

Sound · Computer Science 2023-05-29 Xipin Wei , Junhui Chen , Zirui Zheng , Li Guo , Lantian Li , Dong Wang

We present a unified framework for automatic multitrack music arrangement that enables a single pre-trained symbolic music model to handle diverse arrangement scenarios, including reinterpretation, simplification, and additive generation.…

Sound · Computer Science 2025-11-06 Longshen Ou , Jingwei Zhao , Ziyu Wang , Gus Xia , Qihao Liang , Torin Hopkins Ye Wang

Recent advancements in music source separation have significantly progressed, particularly in isolating vocals, drums, and bass elements from mixed tracks. These developments owe much to the creation and use of large-scale, multitrack…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-18 Jaime Garcia-Martinez , David Diaz-Guerra , Archontis Politis , Tuomas Virtanen , Julio J. Carabias-Orti , Pedro Vera-Candeas

We describe a proof-of-principle implementation of a system for drawing melodies that abstracts away from a note-level input representation via melodic contours. The aim is to allow users to express their musical intentions without…

Sound · Computer Science 2023-05-22 Tashi Namgyal , Peter Flach , Raul Santos-Rodriguez

Video composition is the core task of video editing. Although image composition based on diffusion models has been highly successful, it is not straightforward to extend the achievement to video object composition tasks, which not only…

Computer Vision and Pattern Recognition · Computer Science 2024-06-25 Wei Wang , Yaosen Chen , Yuegen Liu , Qi Yuan , Shubin Yang , Yanru Zhang

Deep generative models have various content creation applications such as graphic design, e-commerce, and virtual Try-on. However, current works mainly focus on synthesizing realistic visual outputs, often ignoring other sensory modalities,…

Computer Vision and Pattern Recognition · Computer Science 2023-05-05 Ruihan Gao , Wenzhen Yuan , Jun-Yan Zhu

Leveraging text, images, structure maps, or motion trajectories as conditional guidance, diffusion models have achieved great success in automated and high-quality video generation. However, generating smooth and rational transition videos…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Zuhao Yang , Jiahui Zhang , Yingchen Yu , Shijian Lu , Song Bai

Music editing has emerged as an important and practical area of artificial intelligence, with applications ranging from video game and film music production to personalizing existing tracks according to user preferences. However, existing…

Sound · Computer Science 2025-11-19 Ali Boudaghi , Hadi Zare

Generating music has a few notable differences from generating images and videos. First, music is an art of time, necessitating a temporal model. Second, music is usually composed of multiple instruments/tracks with their own temporal…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-06 Hao-Wen Dong , Wen-Yi Hsiao , Li-Chia Yang , Yi-Hsuan Yang

Diffusion Transformers (DiTs) have demonstrated remarkable scalability and quality in image and video generation, prompting growing interest in extending them to controllable generation and editing tasks. However, compared to the image…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Ruonan Yu , Zhenxiong Tan , Zigeng Chen , Songhua Liu , Xinchao Wang

Generating music with deep neural networks has been an area of active research in recent years. While the quality of generated samples has been steadily increasing, most methods are only able to exert minimal control over the generated…

Sound · Computer Science 2024-02-23 Dimitri von Rütte , Luca Biggio , Yannic Kilcher , Thomas Hofmann