English
Related papers

Related papers: Exploring Multi-Modal Control in Music-Driven Danc…

200 papers

This study addresses the challenge that generative models struggle to balance flexibility, stability, and controllability in complex interactive scenarios. It proposes a controllable generation framework for dynamic interactive content…

Human-Computer Interaction · Computer Science 2026-02-27 Rui Liu

Effective and generalizable control in video generation remains a significant challenge. While many methods rely on ambiguous or task-specific signals, we argue that a fundamental disentanglement of "appearance" and "motion" provides a more…

Computer Vision and Pattern Recognition · Computer Science 2026-02-16 Mingzhi Sheng , Zekai Gu , Peng Li , Cheng Lin , Hao-Xiang Guo , Ying-Cong Chen , Yuan Liu

Multimodal music creation requires models that can both generate audio from high-level cues and edit existing mixtures in a targeted manner. Yet most multimodal music systems are built for a single task and a fixed prompting interface,…

With the rise of online dance-video platforms and rapid advances in AI-generated content (AIGC), music-driven dance generation has emerged as a compelling research direction. Despite substantial progress in related domains such as…

Computer Vision and Pattern Recognition · Computer Science 2026-05-08 Kaixing Yang , Jiashu Zhu , Xulong Tang , Ziqiao Peng , Xiangyue Zhang , Puwei Wang , Jiahong Wu , Xiangxiang Chu , Hongyan Liu , Jun He

We propose the Multi-Track Music Machine (MMM), a generative system based on the Transformer architecture that is capable of generating multi-track music. In contrast to previous work, which represents musical material as a single…

Sound · Computer Science 2020-08-24 Jeff Ens , Philippe Pasquier

Dancing to music is an instinctive move by humans. Learning to model the music-to-dance generation process is, however, a challenging problem. It requires significant efforts to measure the correlation between music and dance as one needs…

Computer Vision and Pattern Recognition · Computer Science 2019-11-06 Hsin-Ying Lee , Xiaodong Yang , Ming-Yu Liu , Ting-Chun Wang , Yu-Ding Lu , Ming-Hsuan Yang , Jan Kautz

Generating long-term, coherent, and realistic music-conditioned dance sequences remains a challenging task in human motion synthesis. Existing approaches exhibit critical limitations: motion graph methods rely on fixed template libraries,…

Sound · Computer Science 2025-06-04 Mingyang Huang , Peng Zhang , Bang Zhang

Generating 3D models lies at the core of computer graphics and has been the focus of decades of research. With the emergence of advanced neural representations and generative models, the field of 3D content generation is developing rapidly,…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Xiaoyu Li , Qi Zhang , Di Kang , Weihao Cheng , Yiming Gao , Jingbo Zhang , Zhihao Liang , Jing Liao , Yan-Pei Cao , Ying Shan

Character animation in real-world scenarios necessitates a variety of constraints, such as trajectories, key-frames, interactions, etc. Existing methodologies typically treat single or a finite set of these constraint(s) as separate control…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Hanchao Liu , Xiaohang Zhan , Shaoli Huang , Tai-Jiang Mu , Ying Shan

Sound and movement are closely coupled, particularly in dance. Certain audio features have been found to affect the way we move to music. Is this relationship between sound and movement something which can be modelled using machine…

Sound · Computer Science 2020-11-30 Benedikte Wallace , Charles P. Martin , Jim Torresen , Kristian Nymoen

With rapid advances in generative artificial intelligence, the text-to-music synthesis task has emerged as a promising direction for music generation. Nevertheless, achieving precise control over multi-track generation remains an open…

Sound · Computer Science 2024-12-18 Yao Yao , Peike Li , Boyu Chen , Alex Wang

We tackle the task of conditional music generation. We introduce MusicGen, a single Language Model (LM) that operates over several streams of compressed discrete music representation, i.e., tokens. Unlike prior work, MusicGen is comprised…

Creating high-dynamic videos such as motion-rich actions and sophisticated visual effects poses a significant challenge in the field of artificial intelligence. Unfortunately, current state-of-the-art video generation methods, primarily…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Yan Zeng , Guoqiang Wei , Jiani Zheng , Jiaxin Zou , Yang Wei , Yuchen Zhang , Hang Li

The development of generative Machine Learning (ML) models in creative practices, enabled by the recent improvements in usability and availability of pre-trained models, is raising more and more interest among artists, practitioners and…

Machine Learning · Statistics 2022-11-17 Axel Chemla--Romeu-Santos , Philippe Esling

Generative AI has made significant strides in computer vision, particularly in text-driven image/video synthesis (T2I/T2V). Despite the notable advancements, it remains challenging in human-centric content synthesis such as realistic dance…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Tan Wang , Linjie Li , Kevin Lin , Yuanhao Zhai , Chung-Ching Lin , Zhengyuan Yang , Hanwang Zhang , Zicheng Liu , Lijuan Wang

Generating multi-instrument music from symbolic music representations is an important task in Music Information Retrieval (MIR). A central but still largely unsolved problem in this context is musically and acoustically informed control in…

Sound · Computer Science 2023-09-22 Ben Maman , Johannes Zeitler , Meinard Müller , Amit H. Bermano

This paper presents a novel framework for modeling and conditional generation of 3D articulated objects. Troubled by flexibility-quality tradeoffs, existing methods are often limited to using predefined structures or retrieving shapes from…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Jiayi Su , Youhe Feng , Zheng Li , Jinhua Song , Yangfan He , Botao Ren , Botian Xu

Despite the remarkable progress of 3D generation, achieving controllability, i.e., ensuring consistency between generated 3D content and input conditions like edge and depth, remains a significant challenge. Existing methods often struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-08-14 Hongbin Xu , Chaohui Yu , Feng Xiao , Jiazheng Xing , Hai Ci , Weitao Chen , Fan Wang , Ming Li

We demonstrate how conditional generation from diffusion models can be used to tackle a variety of realistic tasks in the production of music in 44.1kHz stereo audio with sampling-time guidance. The scenarios we consider include…

Sound · Computer Science 2023-12-06 Mark Levy , Bruno Di Giorgi , Floris Weers , Angelos Katharopoulos , Tom Nickson

Creation of images using generative adversarial networks has been widely adapted into multi-modal regime with the advent of multi-modal representation models pre-trained on large corpus. Various modalities sharing a common representation…

Sound · Computer Science 2022-06-10 Yoonjeon Kim , Joel Jang , Sumin Shin