English
Related papers

Related papers: RapVerse: Coherent Vocals and Whole-Body Motions G…

200 papers

This work addresses the problem of generating 3D holistic body motions from human speech. Given a speech recording, we synthesize sequences of 3D body poses, hand gestures, and facial expressions that are realistic and diverse. To achieve…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Hongwei Yi , Hualin Liang , Yifei Liu , Qiong Cao , Yandong Wen , Timo Bolkart , Dacheng Tao , Michael J. Black

The field has made significant progress in synthesizing realistic human motion driven by various modalities. Yet, the need for different methods to animate various body parts according to different control signals limits the scalability of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-29 Zixiang Zhou , Yu Wan , Baoyuan Wang

This paper proposes MotionVerse, a unified framework that harnesses the capabilities of Large Language Models (LLMs) to comprehend, generate, and edit human motion in both single-person and multi-person scenarios. To efficiently represent…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Ruibing Hou , Mingshuang Luo , Hongyu Pan , Hong Chang , Shiguang Shan

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

Machine Learning · Computer Science 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow

A song is a combination of singing voice and accompaniment. However, existing works focus on singing voice synthesis and music generation independently. Little attention was paid to explore song synthesis. In this work, we propose a novel…

Audio and Speech Processing · Electrical Eng. & Systems 2024-05-21 Zhiqing Hong , Rongjie Huang , Xize Cheng , Yongqi Wang , Ruiqi Li , Fuming You , Zhou Zhao , Zhimeng Zhang

Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Kunkun Pang , Dafei Qin , Yingruo Fan , Julian Habekost , Takaaki Shiratori , Junichi Yamagishi , Taku Komura

We introduce UniMuMo, a unified multimodal model capable of taking arbitrary text, music, and motion data as input conditions to generate outputs across all three modalities. To address the lack of time-synchronized data, we align unpaired…

Sound · Computer Science 2024-10-08 Han Yang , Kun Su , Yutong Zhang , Jiaben Chen , Kaizhi Qian , Gaowen Liu , Chuang Gan

Our research presents a novel motion generation framework designed to produce whole-body motion sequences conditioned on multiple modalities simultaneously, specifically text and audio inputs. Leveraging Vector Quantized Variational…

Computer Vision and Pattern Recognition · Computer Science 2024-09-04 Sohan Anisetty , James Hays

Text-to-song generation, the task of creating vocals and accompaniment from textual inputs, poses significant challenges due to domain complexity and data scarcity. Existing approaches often employ multi-stage generation procedures, leading…

We propose a novel task for generating 3D dance movements that simultaneously incorporate both text and music modalities. Unlike existing works that generate dance movements using a single modality such as music, our goal is to produce…

Computer Vision and Pattern Recognition · Computer Science 2023-10-03 Kehong Gong , Dongze Lian , Heng Chang , Chuan Guo , Zihang Jiang , Xinxin Zuo , Michael Bi Mi , Xinchao Wang

Whole-body multi-modal human motion generation poses two primary challenges: creating an effective motion generation mechanism and integrating various modalities, such as text, speech, and music, into a cohesive framework. Unlike previous…

Computer Vision and Pattern Recognition · Computer Science 2025-10-17 Zhe Li , Weihao Yuan , Weichao Shen , Siyu Zhu , Zilong Dong , Chang Xu

Video generation models have advanced significantly, yet they still struggle to synthesize complex human movements due to the high degrees of freedom in human articulation. This limitation stems from the intrinsic constraints of pixel-only…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Yuxiao Yang , Hualian Sheng , Sijia Cai , Jing Lin , Jiahao Wang , Bing Deng , Junzhe Lu , Haoqian Wang , Jieping Ye

Text-driven human motion generation is an emerging task in animation and humanoid robot design. Existing algorithms directly generate the full sequence which is computationally expensive and prone to errors as it does not pay special…

Computer Vision and Pattern Recognition · Computer Science 2024-05-27 Zichen Geng , Caren Han , Zeeshan Hayder , Jian Liu , Mubarak Shah , Ajmal Mian

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Lele Chen , Guofeng Cui , Celong Liu , Zhong Li , Ziyi Kou , Yi Xu , Chenliang Xu

This work focuses on full-body co-speech gesture generation. Existing methods typically employ an autoregressive model accompanied by vector-quantized tokens for gesture generation, which results in information loss and compromises the…

Graphics · Computer Science 2025-03-19 Binjie Liu , Lina Liu , Sanyi Zhang , Songen Gu , Yihao Zhi , Tianyi Zhu , Lei Yang , Long Ye

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Shuling Zhao , Fa-Ting Hong , Xiaoshui Huang , Dan Xu

We present a wav-to-wav generative model for the task of singing voice conversion from any identity. Our method utilizes both an acoustic model, trained for the task of automatic speech recognition, together with melody extracted features…

Audio and Speech Processing · Electrical Eng. & Systems 2020-08-10 Adam Polyak , Lior Wolf , Yossi Adi , Yaniv Taigman

Human communication is inherently multimodal, involving a combination of verbal and non-verbal cues such as speech, facial expressions, and body gestures. Modeling these behaviors is essential for understanding human interaction and for…

Computer Vision and Pattern Recognition · Computer Science 2024-12-17 Changan Chen , Juze Zhang , Shrinidhi K. Lakshmikanth , Yusu Fang , Ruizhi Shao , Gordon Wetzstein , Li Fei-Fei , Ehsan Adeli

Well-coordinated, music-aligned holistic dance enhances emotional expressiveness and audience engagement. However, generating such dances remains challenging due to the scarcity of holistic 3D dance datasets, the difficulty of achieving…

Multimedia · Computer Science 2025-07-30 Xiaojie Li , Ronghui Li , Shukai Fang , Shuzhao Xie , Xiaoyang Guo , Jiaqing Zhou , Junkun Peng , Zhi Wang

Inspired by the strong ties between vision and language, the two intimate human sensing and communication modalities, our paper aims to explore the generation of 3D human full-body motions from texts, as well as its reciprocal task,…

Computer Vision and Pattern Recognition · Computer Science 2022-08-08 Chuan Guo , Xinxin Zuo , Sen Wang , Li Cheng
‹ Prev 1 2 3 10 Next ›