English
Related papers

Related papers: MoMu-Diffusion: On Learning Long-Term Motion-Music…

200 papers

Music editing primarily entails the modification of instrument tracks or remixing in the whole, which offers a novel reinterpretation of the original piece through a series of operations. These music processing methods hold immense…

Sound · Computer Science 2023-12-13 Bing Han , Junyu Dai , Weituo Hao , Xinyan He , Dong Guo , Jitong Chen , Yuxuan Wang , Yanmin Qian , Xuchen Song

Text-to-video diffusion models are notoriously limited in their ability to model temporal aspects such as motion, physics, and dynamic interactions. Existing approaches address this limitation by retraining the model or introducing external…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Ariel Shaulov , Itay Hazan , Lior Wolf , Hila Chefer

Diffusion models, particularly latent diffusion models, have demonstrated remarkable success in text-driven human motion generation. However, it remains challenging for latent diffusion models to effectively compose multiple semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-06-05 Jianrong Zhang , Hehe Fan , Yi Yang

We introduce a novel Stylized Motion Diffusion model, dubbed SMooDi, to generate stylized motion driven by content texts and style motion sequences. Unlike existing methods that either generate motion of various content or transfer style…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Lei Zhong , Yiming Xie , Varun Jampani , Deqing Sun , Huaizu Jiang

Recent advances in video diffusion models have unlocked new potential for realistic audio-driven talking video generation. However, achieving seamless audio-lip synchronization, maintaining long-term identity consistency, and producing…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Longtao Zheng , Yifan Zhang , Hanzhong Guo , Jiachun Pan , Zhenxiong Tan , Jiahao Lu , Chuanxin Tang , Bo An , Shuicheng Yan

We present a framework for real-time human-AI musical co-performance, in which a latent diffusion model generates instrumental accompaniment in response to a live stream of context audio. The system combines a MAX/MSP front-end-handling…

Sound · Computer Science 2026-04-10 Tornike Karchkhadze , Shlomo Dubnov

Synthesizing realistic human-object interaction motions is a critical problem in VR/AR and human animation. Unlike the commonly studied scenarios involving a single human or hand interacting with one object, we address a more generic…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Wenkun He , Yun Liu , Ruitao Liu , Li Yi

Motion style transfer is a significant research direction in the field of computer vision, enabling virtual digital humans to rapidly switch between different styles of the same motion, thereby significantly enhancing the richness and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-11 Ziyun Qian , Zeyu Xiao , Xingliang Jin , Dingkang Yang , Mingcheng Li , Zhenyi Wu , Dongliang Kou , Peng Zhai , Lihua Zhang

Recently, human motion analysis has experienced great improvement due to inspiring generative models such as the denoising diffusion model and large language model. While the existing approaches mainly focus on generating motions with…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Yiming Wu , Wei Ji , Kecheng Zheng , Zicheng Wang , Dong Xu

Generating music that temporally aligns with video events is challenging for existing text-to-music models, which lack fine-grained temporal control. We introduce V2M-ZERO, a video-to-music generation approach that generates time-aligned…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Yan-Bo Lin , Jonah Casebeer , Long Mai , Aniruddha Mahapatra , Gedas Bertasius , Nicholas J. Bryan

We introduce Lumina-DiMOO, an open-source foundational model for seamless multi-modal generation and understanding. Lumina-DiMOO sets itself apart from prior unified models by utilizing a fully discrete diffusion modeling to handle inputs…

We introduce a multi-modal diffusion model tailored for the bi-directional conditional generation of video and audio. We propose a joint contrastive training loss to improve the synchronization between visual and auditory occurrences. We…

Machine Learning · Computer Science 2024-10-10 Ruihan Yang , Hannes Gamper , Sebastian Braun

Long-term human motion can be represented as a series of motion modes---motion sequences that capture short-term temporal dynamics---with transitions between them. We leverage this structure and present a novel Motion Transformation…

Machine Learning · Computer Science 2018-08-15 Xinchen Yan , Akash Rastogi , Ruben Villegas , Kalyan Sunkavalli , Eli Shechtman , Sunil Hadap , Ersin Yumer , Honglak Lee

Image animation has become a promising area in multimodal research, with a focus on generating videos from reference images. While prior work has largely emphasized generic video generation guided by text, music-driven dance video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Zhikang Dong , Weituo Hao , Ju-Chiang Wang , Peng Zhang , Pawel Polak

Controllable generation of 3D human motions becomes an important topic as the world embraces digital transformation. Existing works, though making promising progress with the advent of diffusion models, heavily rely on meticulously captured…

Computer Vision and Pattern Recognition · Computer Science 2024-01-25 Nhat M. Hoang , Kehong Gong , Chuan Guo , Michael Bi Mi

Many applications of cross-modal music retrieval are related to connecting sheet music images to audio recordings. A typical and recent approach to this is to learn, via deep neural networks, a joint embedding space that correlates short…

Sound · Computer Science 2023-09-22 Luis Carvalho , Gerhard Widmer

Audio-driven bimanual piano motion generation requires precise modeling of complex musical structures and dynamic cross-hand coordination. However, existing methods often rely on acoustic-only representations lacking symbolic priors, employ…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Xuan Wang , Kai Ruan , Jiayi Han , Kaiyue Zhou , Gaoang Wang

In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based…

Sound · Computer Science 2025-02-28 Xinran Liu , Zhenhua Feng , Diptesh Kanojia , Wenwu Wang

With the introduction of diffusion-based video generation techniques, audio-conditioned human video generation has recently achieved significant breakthroughs in both the naturalness of motion and the synthesis of portrait details. Due to…

Computer Vision and Pattern Recognition · Computer Science 2025-04-07 Jianwen Jiang , Chao Liang , Jiaqi Yang , Gaojie Lin , Tianyun Zhong , Yanbo Zheng

Music is a universal language that can communicate emotions and feelings. It forms an essential part of the whole spectrum of creative media, ranging from movies to social media posts. Machine learning models that can synthesize music are…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Sanjoy Chowdhury , Sayan Nag , K J Joseph , Balaji Vasan Srinivasan , Dinesh Manocha
‹ Prev 1 3 4 5 6 7 10 Next ›