English
Related papers

Related papers: Multimodal Music Generation with Explicit Bridges …

200 papers

This paper explores a specific sub-task of cross-modal music retrieval. We consider the delicate task of retrieving a performance or rendition of a musical piece based on a description of its style, expressive character, or emotion from a…

Sound · Computer Science 2024-01-29 Shreyan Chowdhury , Gerhard Widmer

Every artist has a creative process that draws inspiration from previous artists and their works. Today, "inspiration" has been automated by generative music models. The black box nature of these models obscures the identity of the works…

Sound · Computer Science 2024-01-29 Julia Barnett , Hugo Flores Garcia , Bryan Pardo

Music-to-visual style transfer is a challenging yet important cross-modal learning problem in the practice of creativity. Its major difference from the traditional image style transfer problem is that the style information is provided by…

Computer Vision and Pattern Recognition · Computer Science 2020-09-18 Cheng-Che Lee , Wan-Yi Lin , Yen-Ting Shih , Pei-Yi Patricia Kuo , Li Su

Multi-modal embeddings form the foundation for vision-language models, such as CLIP embeddings, the most widely used text-image embeddings. However, these embeddings are vulnerable to subtle misalignment of cross-modal features, resulting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Yilin Ye , Shishi Xiao , Xingchen Zeng , Wei Zeng

When editing a video, a piece of attractive background music is indispensable. However, video background music generation tasks face several challenges, for example, the lack of suitable training datasets, and the difficulties in flexibly…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Sizhe Li , Yiming Qin , Minghang Zheng , Xin Jin , Yang Liu

Many music AI models learn a map between music content and human-defined labels. However, many annotations, such as chords, can be naturally expressed within the music modality itself, e.g., as sequences of symbolic notes. This observation…

Sound · Computer Science 2025-09-30 Junyan Jiang , Daniel Chin , Liwei Lin , Xuanjie Liu , Gus Xia

This study introduces a text-conditioned approach to generating drumbeats with Latent Diffusion Models (LDMs). It uses informative conditioning text extracted from training data filenames. By pretraining a text and drumbeat encoder through…

Sound · Computer Science 2024-08-07 Pushkar Jajoria , James McDermott

Textual descriptions for multimodal inputs entail recurrent refinement of queries to produce relevant output images. Despite efforts to address challenges such as scaling model size and data volume, the cost associated with pre-training and…

Machine Learning · Computer Science 2025-08-14 Amit Kumar Jaiswal , Haiming Liu , Ingo Frommholz

This thesis combines audio-analysis with computer vision to approach Music Information Retrieval (MIR) tasks from a multi-modal perspective. This thesis focuses on the information provided by the visual layer of music videos and how it can…

Multimedia · Computer Science 2020-02-04 Alexander Schindler

Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first framework to scale…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Shubo Lin , Xuanyang Zhang , Wei Cheng , Weiming Hu , Gang Yu , Jin Gao

Creative generation is the synthesis of new, surprising, and valuable samples that reflect user intent yet cannot be envisioned in advance. This task aims to extend human imagination, enabling the discovery of visual concepts that exist in…

Graphics · Computer Science 2025-10-14 Shelly Golan , Yotam Nitzan , Zongze Wu , Or Patashnik

Conditional music generation offers significant advantages in terms of user convenience and control, presenting great potential in AI-generated content research. However, building conditional generative systems for multitrack popular songs…

Sound · Computer Science 2025-10-27 Jing Luo , Xinyu Yang , Dorien Herremans

Unified audio-visual generation is rapidly gaining industrial and creative relevance, enabling applications in virtual production and interactive media. However, when moving from general audio-video synthesis to music-dance co-generation,…

Solo piano music, despite being a single-instrument medium, possesses significant expressive capabilities, conveying rich semantic information across genres, moods, and styles. However, current general-purpose music representation models,…

Sound · Computer Science 2025-09-05 Hayeon Bang , Eunjin Choi , Seungheon Doh , Juhan Nam

Multimodal Recommender Systems aim to improve recommendation accuracy by integrating heterogeneous content, such as images and textual metadata. While effective, it remains unclear whether their gains stem from true multimodal understanding…

Information Retrieval · Computer Science 2025-08-07 Claudio Pomo , Matteo Attimonelli , Danilo Danese , Fedelucio Narducci , Tommaso Di Noia

We propose a novel approach for the generation of polyphonic music based on LSTMs. We generate music in two steps. First, a chord LSTM predicts a chord progression based on a chord embedding. A second LSTM then generates polyphonic music…

Sound · Computer Science 2017-11-22 Gino Brunner , Yuyi Wang , Roger Wattenhofer , Jonas Wiesendanger

Existing text-to-video diffusion models rely solely on text-only encoders for their pretraining. This limitation stems from the absence of large-scale multimodal prompt video datasets, resulting in a lack of visual grounding and restricting…

Computer Vision and Pattern Recognition · Computer Science 2024-07-10 Yuwei Fang , Willi Menapace , Aliaksandr Siarohin , Tsai-Shien Chen , Kuan-Chien Wang , Ivan Skorokhodov , Graham Neubig , Sergey Tulyakov

We introduce MusicLM, a model generating high-fidelity music from text descriptions such as "a calming violin melody backed by a distorted guitar riff". MusicLM casts the process of conditional music generation as a hierarchical…

Recent advances in video-to-audio (V2A) generation enable high-quality audio synthesis from visual content, yet achieving robust and fine-grained controllability remains challenging. Existing methods suffer from weak textual controllability…

The seamless integration of music with dance movements is essential for communicating the artistic intent of a dance piece. This alignment also significantly improves the immersive quality of gaming experiences and animation productions.…

Sound · Computer Science 2024-09-16 Sifei Li , Weiming Dong , Yuxin Zhang , Fan Tang , Chongyang Ma , Oliver Deussen , Tong-Yee Lee , Changsheng Xu