English
Related papers

Related papers: Video-based Music Generation

200 papers

In this paper, we introduce a MusIc conditioned 3D Dance GEneraTion model, named MIDGET based on Dance motion Vector Quantised Variational AutoEncoder (VQ-VAE) model and Motion Generative Pre-Training (GPT) model to generate vibrant and…

Sound · Computer Science 2024-04-19 Jinwu Wang , Wei Mao , Miaomiao Liu

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Lele Chen , Guofeng Cui , Celong Liu , Zhong Li , Ziyi Kou , Yi Xu , Chenliang Xu

Song generation focuses on producing controllable high-quality songs based on various prompts. However, existing methods struggle to generate vocals and accompaniments with prompt-based control and proper alignment. Additionally, they fall…

Audio and Speech Processing · Electrical Eng. & Systems 2026-02-17 Yu Zhang , Wenxiang Guo , Changhao Pan , Zhiyuan Zhu , Ruiqi Li , Jingyu Lu , Rongjie Huang , Ruiyuan Zhang , Zhiqing Hong , Ziyue Jiang , Zhou Zhao

Recent advances in audio-synchronized visual animation enable control of video content using audios from specific classes. However, existing methods rely heavily on expensive manual curation of high-quality, class-specific training videos,…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Lin Zhang , Zefan Cai , Yufan Zhou , Shentong Mo , Jinhong Lin , Cheng-En Wu , Yibing Wei , Yijing Zhang , Ruiyi Zhang , Wen Xiao , Tong Sun , Junjie Hu , Pedro Morgado

We present DanceAnyWay, a generative learning method to synthesize beat-guided dances of 3D human characters synchronized with music. Our method learns to disentangle the dance movements at the beat frames from the dance movements at all…

Sound · Computer Science 2024-11-26 Aneesh Bhattacharya , Manas Paranjape , Uttaran Bhattacharya , Aniket Bera

Recent years have witnessed significant progress in generative models for music, featuring diverse architectures that balance output quality, diversity, speed, and user control. This study explores a user-friendly graphical interface…

Sound · Computer Science 2024-07-02 Scott H. Hawley

In this work, we introduce Mozualization, a music generation and editing tool that creates multi-style embedded music by integrating diverse inputs, such as keywords, images, and sound clips (e.g., segments from various pieces of music or…

Human-Computer Interaction · Computer Science 2025-04-22 Wanfang Xu , Lixiang Zhao , Haiwen Song , Xinheng Song , Zhaolin Lu , Yu Liu , Min Chen , Eng Gee Lim , Lingyun Yu

We present the MIDInfinite, a web application capable of generating symbolic music using a large-scale generative AI model locally on commodity hardware. Creating this demo involved porting the Anticipatory Music Transformer, a large…

Sound · Computer Science 2024-11-15 Xun Zhou , Charlie Ruan , Zihe Zhao , Tianqi Chen , Chris Donahue

High-quality synthetic data can support the development of effective predictive models for biomedical tasks, especially in rare diseases or when subject to compelling privacy constraints. These limitations, for instance, negatively impact…

Machine Learning · Computer Science 2023-01-24 Lorenzo Simone , Davide Bacciu

Recent studies in singing voice synthesis have achieved high-quality results leveraging advances in text-to-speech models based on deep neural networks. One of the main issues in training singing voice synthesis models is that they require…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-15 Soonbeom Choi , Juhan Nam

We present and release MIDI-GPT, a generative system based on the Transformer architecture that is designed for computer-assisted music composition workflows. MIDI-GPT supports the infilling of musical material at the track and bar level,…

State-of-the-art symbolic music generation models have recently achieved remarkable output quality, yet explicit control over compositional features, such as tonal tension, remains challenging. We propose a novel approach that integrates a…

Sound · Computer Science 2025-11-25 Maral Ebrahimzadeh , Gilberto Bernardes , Sebastian Stober

Both images and music can convey rich semantics and are widely used to induce specific emotions. Matching images and music with similar emotions might help to make emotion perceptions more vivid and stronger. Existing emotion-based image…

Computer Vision and Pattern Recognition · Computer Science 2020-09-14 Sicheng Zhao , Yaxian Li , Xingxu Yao , Weizhi Nie , Pengfei Xu , Jufeng Yang , Kurt Keutzer

Creativity, or the ability to produce new useful ideas, is commonly associated to the human being; but there are many other examples in nature where this phenomenon can be observed. Inspired by this fact, in engineering and particularly in…

Sound · Computer Science 2022-01-26 David Daniel Albarracín Molina

We introduce EQUI-VOCAL: a new system that automatically synthesizes queries over videos from limited user interactions. The user only provides a handful of positive and negative examples of what they are looking for. EQUI-VOCAL utilizes…

Databases · Computer Science 2023-08-09 Enhao Zhang , Maureen Daum , Dong He , Brandon Haynes , Ranjay Krishna , Magdalena Balazinska

The ability to automatically generate music that appropriately matches an arbitrary input track is a challenging task. We present a novel controllable system for generating single stems to accompany musical mixes of arbitrary length. At the…

Sound · Computer Science 2024-02-05 Marco Pasini , Maarten Grachten , Stefan Lattner

Emotional information is essential for enhancing human-computer interaction and deepening image understanding. However, while deep learning has advanced image recognition, the intuitive understanding and precise control of emotional…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Junjie Xu , Xingjiao Wu , Tanren Yao , Zihao Zhang , Jiayang Bei , Wu Wen , Liang He

We propose a novel system that takes as an input body movements of a musician playing a musical instrument and generates music in an unsupervised setting. Learning to generate multi-instrumental music from videos without labeling the…

Sound · Computer Science 2020-12-08 Kun Su , Xiulong Liu , Eli Shlizerman

Lyric-to-melody generation is a highly challenging task in the field of AI music generation. Due to the difficulty of learning strict yet weak correlations between lyrics and melodies, previous methods have suffered from weak…

Audio and Speech Processing · Electrical Eng. & Systems 2025-01-16 Li Chai , Donglin Wang

Synthesize human motions from music, i.e., music to dance, is appealing and attracts lots of research interests in recent years. It is challenging due to not only the requirement of realistic and complex human motions for dance, but more…

Computer Vision and Pattern Recognition · Computer Science 2020-03-12 Wenlin Zhuang , Congyi Wang , Siyu Xia , Jinxiang Chai , Yangang Wang
‹ Prev 1 8 9 10 Next ›