English
Related papers

Related papers: Mozart's Touch: A Lightweight Multi-modal Music Ge…

200 papers

Music generation introduces challenging complexities to large language models. Symbolic structures of music often include vertical harmonization as well as horizontal counterpoint, urging various adaptations and enhancements for large-scale…

Sound · Computer Science 2024-07-30 Seungyeon Rhyu , Kichang Yang , Sungjun Cho , Jaehyeon Kim , Kyogu Lee , Moontae Lee

Multimodal large language models (MLLMs) extend the success of language models to visual understanding, and recent efforts have sought to build unified MLLMs that support both understanding and generation. However, constructing such models…

Computer Vision and Pattern Recognition · Computer Science 2025-10-03 Hanyu Wang , Jiaming Han , Ziyan Yang , Qi Zhao , Shanchuan Lin , Xiangyu Yue , Abhinav Shrivastava , Zhenheng Yang , Hao Chen

Musical mode is one of the most critical element that establishes the framework of pitch organization and determines the harmonic relationships. Previous works often use the simplistic and rigid alignment method, and overlook the diversity…

Sound · Computer Science 2025-01-15 Qian Liang , Yi Zeng , Menghaoran Tang

Existing music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating…

Sound · Computer Science 2024-03-21 Ronghui Li , Yuqin Dai , Yachao Zhang , Jun Li , Jian Yang , Jie Guo , Xiu Li

Multimodal Generative Models (MGMs) have rapidly evolved beyond text generation, now spanning diverse output modalities including images, music, video, human motion, and 3D objects, by integrating language with other sensory modalities…

Multimedia · Computer Science 2025-11-25 Longzhen Han , Awes Mubarak , Almas Baimagambetov , Nikolaos Polatidis , Thar Baker

Lyrics-to-melody generation is an interesting and challenging topic in AI music research field. Due to the difficulty of learning the correlations between lyrics and melody, previous methods suffer from low generation quality and lack of…

Sound · Computer Science 2023-06-06 Zhe Zhang , Yi Yu , Atsuhiro Takasu

In the task of generating music, the art factor plays a big role and is a great challenge for AI. Previous work involving adversarial training to produce new music pieces and modeling the compatibility of variety in music (beats, tempo,…

Sound · Computer Science 2023-01-09 Abhinav Kaushal Keshari

Developing text-driven symbolic music generation models remains challenging due to the scarcity of aligned text-music datasets and the unreliability of automated captioning pipelines. While most efforts have focused on MIDI, sheet music…

The development of generative Machine Learning (ML) models in creative practices, enabled by the recent improvements in usability and availability of pre-trained models, is raising more and more interest among artists, practitioners and…

Machine Learning · Statistics 2022-11-17 Axel Chemla--Romeu-Santos , Philippe Esling

In recent years, the use of large language models (LLMs) to generate music content, particularly lyrics, has gained in popularity. These advances provide valuable tools for artists and enhance their creative processes, but they also raise…

Computation and Language · Computer Science 2025-04-25 Yanis Labrak , Markus Frohmann , Gabriel Meseguer-Brocal , Elena V. Epure

We have recently seen tremendous progress in realistic text-to-motion generation. Yet, the existing methods often fail or produce implausible motions with unseen text inputs, which limits the applications. In this paper, we present OMG, a…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Han Liang , Jiacheng Bao , Ruichi Zhang , Sihan Ren , Yuecheng Xu , Sibei Yang , Xin Chen , Jingyi Yu , Lan Xu

AI-based music generation has made significant progress in recent years. However, generating symbolic music that is both long-structured and expressive remains a significant challenge. In this paper, we propose PerceiverS (Segmentation and…

Artificial Intelligence · Computer Science 2025-09-23 Yungang Yi , Weihua Li , Matthew Kuo , Quan Bai

Audio is inherently temporal and closely synchronized with the visual world, making it a naturally aligned and expressive control signal for controllable video generation (e.g., movies). Beyond control, directly translating audio into video…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Shuchen Weng , Haojie Zheng , Zheng Chang , Si Li , Boxin Shi , Xinlong Wang

This study proposes a system designed to enumerate the process of collaborative composition among humans, using automatic music composition technology. By integrating multiple Recurrent Neural Network (RNN) models, the system provides an…

Sound · Computer Science 2024-03-07 So Hirawata , Noriko Otani

Expressive synthetic speech is essential for many human-computer interaction and audio broadcast scenarios, and thus synthesizing expressive speech has attracted much attention in recent years. Previous methods performed the expressive…

Sound · Computer Science 2022-01-19 Yi Lei , Shan Yang , Xinsheng Wang , Lei Xie

Recently, emotional talking face generation has received considerable attention. However, existing methods only adopt one-hot coding, image, or audio as emotion conditions, thus lacking flexible control in practical applications and failing…

Computer Vision and Pattern Recognition · Computer Science 2023-06-01 Chao Xu , Junwei Zhu , Jiangning Zhang , Yue Han , Wenqing Chu , Ying Tai , Chengjie Wang , Zhifeng Xie , Yong Liu

Synthesizing human motion with a global structure, such as a choreography, is a challenging task. Existing methods tend to concentrate on local smooth pose transitions and neglect the global context or the theme of the motion. In this work,…

In this work, we study the problem of generating novel images from complex multimodal prompt sequences. While existing methods achieve promising results for text-to-image generation, they often struggle to capture fine-grained details from…

Computer Vision and Pattern Recognition · Computer Science 2024-05-29 Amandeep Kumar , Muzammal Naseer , Sanath Narayan , Rao Muhammad Anwer , Salman Khan , Hisham Cholakkal

The growing diversity of large language models (LLMs) means users often need to compare and combine outputs from different models to obtain higher-quality or more comprehensive responses. However, switching between separate interfaces and…

Human-Computer Interaction · Computer Science 2025-10-23 Yingtian Shi , Jinda Yang , Yuhan Wang , Yiwen Yin , Haoyu Li , Kunyu Gao , Chun Yu

Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language generation and understanding remains largely underexplored.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Zekun Li , Sizhe An , Chengcheng Tang , Chuan Guo , Ivan Shugurov , Linguang Zhang , Amy Zhao , Srinath Sridhar , Lingling Tao , Abhay Mittal