English
Related papers

Related papers: Music- and Lyrics-driven Dance Synthesis

200 papers

We present Dance2Music-GAN (D2M-GAN), a novel adversarial multi-modal framework that generates complex musical samples conditioned on dance videos. Our proposed framework takes dance video frames and human body motions as input, and learns…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Ye Zhu , Kyle Olszewski , Yu Wu , Panos Achlioptas , Menglei Chai , Yan Yan , Sergey Tulyakov

This dissertation proposes the study of multimodal learning in the context of musical signals. Throughout, we focus on the interaction between audio signals and text information. Among the many text sources related to music that can be used…

Sound · Computer Science 2021-11-01 Gabriel Meseguer-Brocal

Automatic melody-to-lyric generation is a task in which song lyrics are generated to go with a given melody. It is of significant practical interest and more challenging than unconstrained lyric generation as the music imposes additional…

Text-to-motion generation holds significant potential for cross-linguistic applications, yet it is hindered by the lack of bilingual datasets and the poor cross-lingual semantic understanding of existing language models. To address these…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Wanjiang Weng , Xiaofeng Tan , Xiangbo Shu , Guo-Sen Xie , Pan Zhou , Hongsong Wang

Human motion generation has advanced rapidly in recent years, yet the critical problem of creating spatially grounded, context-aware gestures has been largely overlooked. Existing models typically specialize either in descriptive motion…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Anna Deichler , Jim O'Regan , Teo Guichoux , David Johansson , Jonas Beskow

Linking human motion and natural language is of great interest for the generation of semantic representations of human activities as well as for the generation of robot activities based on natural language input. However, while there have…

Robotics · Computer Science 2018-08-10 Matthias Plappert , Christian Mandery , Tamim Asfour

In this work, we propose TextIM, a novel framework for synthesizing TEXT-driven human Interactive Motions, with a focus on the precise alignment of part-level semantics. Existing methods often overlook the critical roles of interactive body…

Computer Vision and Pattern Recognition · Computer Science 2024-08-07 Siyuan Fan , Bo Du , Xiantao Cai , Bo Peng , Longling Sun

Text-guided 3D motion editing has seen success in single-person scenarios, but its extension to multi-person settings is less explored due to limited paired data and the complexity of inter-person interactions. We introduce the task of…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Yebin Yang , Di Wen , Lei Qi , Weitong Kong , Junwei Zheng , Ruiping Liu , Yufan Chen , Chengzhi Wu , Kailun Yang , Yuqian Fu , Danda Pani Paudel , Luc Van Gool , Kunyu Peng

At the core of many important machine learning problems faced by online streaming services is a need to model how users interact with the content they are served. Unfortunately, there are no public datasets currently available that enable…

Information Retrieval · Computer Science 2020-10-16 Brian Brost , Rishabh Mehrotra , Tristan Jehan

We address the challenging problem of fine-grained text-driven human motion generation. Existing works generate imprecise motions that fail to accurately capture relationships specified in text due to: (1) lack of effective text parsing for…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Yin Wang , Mu Li , Jiapeng Liu , Zhiying Leng , Frederick W. B. Li , Ziyao Zhang , Xiaohui Liang

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yunnan Wang , Kecheng Zheng , Jianyuan Wang , Minghao Chen , David Novotny , Christian Rupprecht , Yinghao Xu , Xing Zhu , Wenjun Zeng , Xin Jin , Yujun Shen

Dance videos are interesting and semantics-intensive. At the same time, they are the complex type of videos compared to all other types such as sports, news and movie videos. In fact, dance video is the one which is less explored by the…

Multimedia · Computer Science 2010-01-05 Rajkumar Kannan , Balakrishnan Ramadoss

With the growing demands of AI-generated content (AIGC), the need for high-quality, diverse, and scalable data has become increasingly crucial. However, collecting large-scale real-world data remains costly and time-consuming, hindering the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-08 Kunyu Feng , Yue Ma , Xinhua Zhang , Boshi Liu , Yikuang Yuluo , Yinhan Zhang , Runtao Liu , Hongyu Liu , Zhiyuan Qin , Shanhui Mo , Qifeng Chen , Zeyu Wang

Memes are the new-age conveyance mechanism for humor on social media sites. Memes often include an image and some text. Memes can be used to promote disinformation or hatred, thus it is crucial to investigate in details. We introduce…

Human motion synthesis is a fundamental task in computer animation. Despite recent progress in this field utilizing deep learning and motion capture data, existing methods are always limited to specific motion categories, environments, and…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Zhikai Zhang , Yitang Li , Haofeng Huang , Mingxian Lin , Li Yi

Text-conditioned video diffusion models have emerged as a powerful tool in the realm of video generation and editing. But their ability to capture the nuances of human movement remains under-explored. Indeed the ability of these models to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Paul Janson , Tiberiu Popa , Eugene Belilovsky

Recent advancements in music source separation have significantly progressed, particularly in isolating vocals, drums, and bass elements from mixed tracks. These developments owe much to the creation and use of large-scale, multitrack…

Audio and Speech Processing · Electrical Eng. & Systems 2025-02-18 Jaime Garcia-Martinez , David Diaz-Guerra , Archontis Politis , Tuomas Virtanen , Julio J. Carabias-Orti , Pedro Vera-Candeas

Automatic song writing is a topic of significant practical interest. However, its research is largely hindered by the lack of training data due to copyright concerns and challenged by its creative nature. Most noticeably, prior works often…

We present Animus3D, a text-driven 3D animation framework that generates motion field given a static 3D asset and text prompt. Previous methods mostly leverage the vanilla Score Distillation Sampling (SDS) objective to distill motion from…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Qi Sun , Can Wang , Jiaxiang Shang , Wensen Feng , Jing Liao

Dance and music are two highly correlated artistic forms. Synthesizing dance motions has attracted much attention recently. Most previous works conduct music-to-dance synthesis via directly music to human skeleton keypoints mapping.…

Computer Vision and Pattern Recognition · Computer Science 2020-09-17 Zijie Ye , Haozhe Wu , Jia Jia , Yaohua Bu , Wei Chen , Fanbo Meng , Yanfeng Wang