English
Related papers

Related papers: DisCo: Disentangled Control for Realistic Human Da…

200 papers

Dance is an important human art form, but creating new dances can be difficult and time-consuming. In this work, we introduce Editable Dance GEneration (EDGE), a state-of-the-art method for editable dance generation that is capable of…

Sound · Computer Science 2022-11-29 Jonathan Tseng , Rodrigo Castellon , C. Karen Liu

Text-video retrieval is a challenging cross-modal task, which aims to align visual entities with natural language descriptions. Current methods either fail to leverage the local details or are computationally expensive. What's worse, they…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Peng Jin , Hao Li , Zesen Cheng , Jinfa Huang , Zhennan Wang , Li Yuan , Chang Liu , Jie Chen

We present Dance2Music-GAN (D2M-GAN), a novel adversarial multi-modal framework that generates complex musical samples conditioned on dance videos. Our proposed framework takes dance video frames and human body motions as input, and learns…

Computer Vision and Pattern Recognition · Computer Science 2022-07-20 Ye Zhu , Kyle Olszewski , Yu Wu , Panos Achlioptas , Menglei Chai , Yan Yan , Sergey Tulyakov

Human motion generation is a significant pursuit in generative computer vision with widespread applications in film-making, video games, AR/VR, and human-robot interaction. Current methods mainly utilize either diffusion-based generative…

Computer Vision and Pattern Recognition · Computer Science 2025-02-03 Canxuan Gang

Dance typically involves professional choreography with complex movements that follow a musical rhythm and can also be influenced by lyrical content. The integration of lyrics in addition to the auditory dimension, enriches the foundational…

Sound · Computer Science 2024-03-15 Wenjie Yin , Xuejiao Zhao , Yi Yu , Hang Yin , Danica Kragic , Mårten Björkman

This paper presents a novel method for generating diverse 3D human poses in scenes with semantic control. Existing methods heavily rely on the human-scene interaction dataset, resulting in a limited diversity of the generated human poses.…

Computer Vision and Pattern Recognition · Computer Science 2024-06-11 Bowen Dang , Xi Zhao

Person image synthesis, e.g., pose transfer, is a challenging problem due to large variation and occlusion. Existing methods have difficulties predicting reasonable invisible regions and fail to decouple the shape and style of clothing,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-01 Jinsong Zhang , Kun Li , Yu-Kun Lai , Jingyu Yang

While recent advancements in text-to-video diffusion models enable high-quality short video generation from a single prompt, generating real-world long videos in a single pass remains challenging due to limited data and high computational…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Subin Kim , Seoung Wug Oh , Jui-Hsien Wang , Joon-Young Lee , Jinwoo Shin

Text-conditioned video diffusion models have emerged as a powerful tool in the realm of video generation and editing. But their ability to capture the nuances of human movement remains under-explored. Indeed the ability of these models to…

Computer Vision and Pattern Recognition · Computer Science 2024-11-21 Paul Janson , Tiberiu Popa , Eugene Belilovsky

In music-driven dance motion generation, most existing methods use hand-crafted features and neglect that music foundation models have profoundly impacted cross-modal content generation. To bridge this gap, we propose a diffusion-based…

Sound · Computer Science 2025-02-28 Xinran Liu , Zhenhua Feng , Diptesh Kanojia , Wenwu Wang

Transferring human motion from a source to a target person poses great potential in computer vision and graphics applications. A crucial step is to manipulate sequential future motion while retaining the appearance characteristic.Previous…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Bowen Wu , Zhenyu Xie , Xiaodan Liang , Yubei Xiao , Haoye Dong , Liang Lin

Choreographers determine what the dances look like, while cameramen determine the final presentation of dances. Recently, various methods and datasets have showcased the feasibility of dance synthesis. However, camera movement synthesis…

Computer Vision and Pattern Recognition · Computer Science 2024-03-21 Zixuan Wang , Jia Jia , Shikun Sun , Haozhe Wu , Rong Han , Zhenyu Li , Di Tang , Jiaqing Zhou , Jiebo Luo

Learning disentangled representations of data is a fundamental problem in artificial intelligence. Specifically, disentangled latent representations allow generative models to control and compose the disentangled factors in the synthesis…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Yotam Nitzan , Amit Bermano , Yangyan Li , Daniel Cohen-Or

Existing music-driven 3D dance generation methods mainly concentrate on high-quality dance generation, but lack sufficient control during the generation process. To address these issues, we propose a unified framework capable of generating…

Sound · Computer Science 2024-03-21 Ronghui Li , Yuqin Dai , Yachao Zhang , Jun Li , Jian Yang , Jie Guo , Xiu Li

The pose-guided person image generation task requires synthesizing photorealistic images of humans in arbitrary poses. The existing approaches use generative adversarial networks that do not necessarily maintain realistic textures or need…

Computer Vision and Pattern Recognition · Computer Science 2023-03-02 Ankan Kumar Bhunia , Salman Khan , Hisham Cholakkal , Rao Muhammad Anwer , Jorma Laaksonen , Mubarak Shah , Fahad Shahbaz Khan

Diffusion models have emerged as the dominant paradigm for style transfer, but their text-driven mechanism is hindered by a core limitation: it treats textual descriptions as uniform, monolithic guidance. This limitation overlooks the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yuanlin Yang , Quanjian Song , Zhexian Gao , Ge Wang , Shanshan Li , Xiaoyan Zhang

Dance-driven music generation aims to generate musical pieces conditioned on dance videos. Previous works focus on monophonic or raw audio generation, while the multi-instruments scenario is under-explored. The challenges associated with…

Multimedia · Computer Science 2024-02-28 Bo Han , Yuheng Li , Yixuan Shen , Yi Ren , Feilin Han

Recent advances in model architectures, compute, and data scale have driven rapid progress in video generation, producing increasingly realistic content. Yet, no prior method systematically measures how faithfully these systems render human…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Yusu Fang , Tiange Xiang , Tian Tan , Narayan Schuetz , Scott Delp , Li Fei-Fei , Ehsan Adeli

Dance serves as a powerful medium for expressing human emotions, but the lifelike generation of dance is still a considerable challenge. Recently, diffusion models have showcased remarkable generative abilities across various domains. They…

Sound · Computer Science 2024-06-25 Canyu Zhang , Youbao Tang , Ning Zhang , Ruei-Sung Lin , Mei Han , Jing Xiao , Song Wang

Synthesizing realistic human-object interactions (HOI) in video is challenging due to the complex, instance-specific interaction dynamics of both humans and objects. Incorporating controllability in video generation further adds to the…

Computer Vision and Pattern Recognition · Computer Science 2026-04-09 Wanyue Zhang , Lin Geng Foo , Thabo Beeler , Rishabh Dabral , Christian Theobalt