English
Related papers

Related papers: DanceCamera3D: 3D Camera Movement Synthesis with M…

200 papers

Generative AI has made significant strides in computer vision, particularly in text-driven image/video synthesis (T2I/T2V). Despite the notable advancements, it remains challenging in human-centric content synthesis such as realistic dance…

Computer Vision and Pattern Recognition · Computer Science 2024-04-08 Tan Wang , Linjie Li , Kevin Lin , Yuanhao Zhai , Chung-Ching Lin , Zhengyuan Yang , Hanwang Zhang , Zicheng Liu , Lijuan Wang

The use of cameras and computational algorithms for noninvasive, low-cost and scalable measurement of physiological (e.g., cardiac and pulmonary) vital signs is very attractive. However, diverse data representing a range of environments,…

Computer Vision and Pattern Recognition · Computer Science 2022-06-10 Daniel McDuff , Miah Wander , Xin Liu , Brian L. Hill , Javier Hernandez , Jonathan Lester , Tadas Baltrusaitis

Dance-driven music generation aims to generate musical pieces conditioned on dance videos. Previous works focus on monophonic or raw audio generation, while the multi-instruments scenario is under-explored. The challenges associated with…

Multimedia · Computer Science 2024-02-28 Bo Han , Yuheng Li , Yixuan Shen , Yi Ren , Feilin Han

Dance serves as both a cultural cornerstone and a medium for personal expression, yet the rapid growth of online dance content has made personalized discovery increasingly difficult. Text-based dance retrieval offers a natural interface for…

Multimedia · Computer Science 2026-05-04 Yawen Qin , Ke Qiu , Qin Zhang

Despite considerable efforts to enhance the generalization of 3D pose estimators without costly 3D annotations, existing data augmentation methods struggle in real world scenarios with diverse human appearances and complex poses. We propose…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 ChangHee Yang , Hyeonseop Song , Seokhun Choi , Seungwoo Lee , Jaechul Kim , Hoseok Do

Generating dance from music is crucial for advancing automated choreography. Current methods typically produce skeleton keypoint sequences instead of dance videos and lack the capability to make specific individuals dance, which reduces…

Computer Vision and Pattern Recognition · Computer Science 2024-12-02 Xuanchen Wang , Heng Wang , Dongnan Liu , Weidong Cai

4D scans of dynamic deformable human body parts help researchers have a better understanding of spatiotemporal features. However, reconstructing 4D scans based on multiple asynchronous cameras encounters two main challenges: 1) finding the…

Image and Video Processing · Electrical Eng. & Systems 2023-07-25 Farzam Tajdari , Toon Huysmans , Xinhe Yao , Jun Xu , Yu Song

The objective of the multi-condition human motion synthesis task is to incorporate diverse conditional inputs, encompassing various forms like text, music, speech, and more. This endows the task with the capability to adapt across multiple…

Computer Vision and Pattern Recognition · Computer Science 2024-04-18 Zeyu Ling , Bo Han , Yongkang Wong , Mohan Kangkanhalli , Weidong Geng

In this paper we address the problem of motion event detection in athlete recordings from individual sports. In contrast to recent end-to-end approaches, we propose to use 2D human pose sequences as an intermediate representation that…

Computer Vision and Pattern Recognition · Computer Science 2020-04-23 Moritz Einfalt , Rainer Lienhart

Technologies play an increasingly important role in sports and become a real competitive advantage for the athletes who benefit from it. Among them, the use of motion capture is developing in various sports to optimize sporting gestures.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-09 Fiche Guénolé , Sevestre Vincent , Gonzalez-Barral Camila , Leglaive Simon , Séguier Renaud

Controllable video generation (CVG) has advanced rapidly, yet current systems falter when more than one actor must move, interact, and exchange positions under noisy control signals. We address this gap with DanceTogether, the first…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Junhao Chen , Mingjin Chen , Jianjin Xu , Xiang Li , Junting Dong , Mingze Sun , Puhua Jiang , Hongxiang Li , Yuhang Yang , Hao Zhao , Xiaoxiao Long , Ruqi Huang

Music-driven 3D dance generation has attracted increasing attention in recent years, with promising applications in choreography, virtual reality, and creative content creation. Previous research has generated promising realistic dance…

Sound · Computer Science 2026-02-24 Kaixing Yang , Xulong Tang , Ziqiao Peng , Yuxuan Hu , Jun He , Hongyan Liu

Generating realistic human videos remains a challenging task, with the most effective methods currently relying on a human motion sequence as a control signal. Existing approaches often use existing motion extracted from other videos, which…

Computer Vision and Pattern Recognition · Computer Science 2024-12-18 Hsin-Ping Huang , Yang Zhou , Jui-Hsien Wang , Difan Liu , Feng Liu , Ming-Hsuan Yang , Zhan Xu

Existing deep models predict 2D and 3D kinematic poses from video that are approximately accurate, but contain visible errors that violate physical constraints, such as feet penetrating the ground and bodies leaning at extreme angles. In…

Computer Vision and Pattern Recognition · Computer Science 2020-07-27 Davis Rempe , Leonidas J. Guibas , Aaron Hertzmann , Bryan Russell , Ruben Villegas , Jimei Yang

In this work, we present DreamDance, a novel method for animating human images using only skeleton pose sequences as conditional inputs. Existing approaches struggle with generating coherent, high-quality content in an efficient and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Yatian Pang , Bin Zhu , Bin Lin , Mingzhe Zheng , Francis E. H. Tay , Ser-Nam Lim , Harry Yang , Li Yuan

Recent works on dynamic 3D neural field reconstruction assume the input from synchronized multi-view videos whose poses are known. The input constraints are often not satisfied in real-world setups, making the approach impractical. We show…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Changwoon Choi , Jeongjun Kim , Geonho Cha , Minkwan Kim , Dongyoon Wee , Young Min Kim

Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D world, depend more on physical properties-such as distance,…

Well-coordinated, music-aligned holistic dance enhances emotional expressiveness and audience engagement. However, generating such dances remains challenging due to the scarcity of holistic 3D dance datasets, the difficulty of achieving…

Multimedia · Computer Science 2025-07-30 Xiaojie Li , Ronghui Li , Shukai Fang , Shuzhao Xie , Xiaoyang Guo , Jiaqing Zhou , Junkun Peng , Zhi Wang

The Codec Avatars Lab at Meta introduces Embody 3D, a multimodal dataset of 500 individual hours of 3D motion data from 439 participants collected in a multi-camera collection stage, amounting to over 54 million frames of tracked 3D motion.…

Recent advances in video diffusion transformers have enabled interactive gaming world models that allow users to explore generated environments over extended horizons. However, existing approaches struggle with precise action control and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-18 Jisu Nam , Yicong Hong , Chun-Hao Paul Huang , Feng Liu , JoungBin Lee , Jiyoung Kim , Siyoon Jin , Yunsung Lee , Jaeyoon Jung , Suhwan Choi , Seungryong Kim , Yang Zhou