English
Related papers

Related papers: Freeform Body Motion Generation from Speech

200 papers

Given an arbitrary face image and an arbitrary speech clip, the proposed work attempts to generating the talking face video with accurate lip synchronization while maintaining smooth transition of both lip and facial movement over the…

Computer Vision and Pattern Recognition · Computer Science 2019-07-29 Yang Song , Jingwen Zhu , Dawei Li , Xiaolong Wang , Hairong Qi

Speech-driven gesture generation aims at synthesizing a gesture sequence synchronized with the input speech signal. Previous methods leverage neural networks to directly map a compact audio representation to the gesture sequence, ignoring…

Computer Vision and Pattern Recognition · Computer Science 2024-10-18 Fengqi Liu , Hexiang Wang , Jingyu Gong , Ran Yi , Qianyu Zhou , Xuequan Lu , Jiangbo Lu , Lizhuang Ma

While there has been significant progress towards modelling coherence in written discourse, the work in modelling spoken discourse coherence has been quite limited. Unlike the coherence in text, coherence in spoken discourse is also…

Computation and Language · Computer Science 2021-01-05 Rajaswa Patil , Yaman Kumar Singla , Rajiv Ratn Shah , Mika Hama , Roger Zimmermann

We present a generative model that learns to synthesize human motion from limited training sequences. Our framework provides conditional generation and blending across multiple temporal resolutions. The model adeptly captures human motion…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 David Eduardo Moreno-Villamarín , Anna Hilsmann , Peter Eisert

Conditional motion generation has been extensively studied in computer vision, yet two critical challenges remain. First, while masked autoregressive methods have recently outperformed diffusion-based approaches, existing masking models…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Zeyu Zhang , Yiran Wang , Wei Mao , Danning Li , Rui Zhao , Biao Wu , Zirui Song , Bohan Zhuang , Ian Reid , Richard Hartley

Generating conversational gestures from speech audio is challenging due to the inherent one-to-many mapping between audio and body motions. Conventional CNNs/RNNs assume one-to-one mapping, and thus tend to predict the average of all…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Jing Li , Di Kang , Wenjie Pei , Xuefei Zhe , Ying Zhang , Zhenyu He , Linchao Bao

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

Computer Vision and Pattern Recognition · Computer Science 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu

Audio-driven human animation methods, such as talking head and talking body generation, have made remarkable progress in generating synchronized facial movements and appealing visual quality videos. However, existing methods primarily focus…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Zhe Kong , Feng Gao , Yong Zhang , Zhuoliang Kang , Xiaoming Wei , Xunliang Cai , Guanying Chen , Wenhan Luo

Speech-driven gesture generation is highly challenging due to the random jitters of human motion. In addition, there is an inherent asynchronous relationship between human speech and gestures. To tackle these challenges, we introduce a…

Human-Computer Interaction · Computer Science 2023-05-19 Sicheng Yang , Zhiyong Wu , Minglei Li , Zhensong Zhang , Lei Hao , Weihong Bao , Haolin Zhuang

Text-driven human motion generation is a multimodal task that synthesizes human motion sequences conditioned on natural language. It requires the model to satisfy textual descriptions under varying conditional inputs, while generating…

Computer Vision and Pattern Recognition · Computer Science 2024-10-01 Xingyu Chen

Human movement studies and analyses have been fundamental in many scientific domains, ranging from neuroscience to education, pattern recognition to robotics, health care to sports, and beyond. Previous speech motor models were proposed to…

Neurons and Cognition · Quantitative Biology 2024-02-01 C. Carmona-Duarte , M. A. Ferrer , R. Plamondon , A. Gomez-Rodellar , P. Gomez-Vilda

Text-driven human motion generation in computer vision is both significant and challenging. However, current methods are limited to producing either deterministic or imprecise motion sequences, failing to effectively control the temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-09-13 Yin Wang , Zhiying Leng , Frederick W. B. Li , Shun-Cheng Wu , Xiaohui Liang

Text-to-motion (T2M) generation has broad applications in character animation, virtual avatars, and human-robot interaction. Existing methods typically generate pose trajectories or motion tokens directly from language, forcing a single…

Machine Learning · Computer Science 2026-05-29 Nikolay Shvetsov , Maksim Bobrin , Nazar Buzun , Dmitry V. Dylov

Sign language video generation requires producing natural signing motions with realistic appearances under precise semantic control, yet faces two critical challenges: excessive signer-specific data requirements and poor generalization. We…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Jiayi He , Xu Wang , Shengeng Tang , Yaxiong Wang , Lechao Cheng , Dan Guo

Co-speech gesture generation is crucial for creating lifelike avatars and enhancing human-computer interactions by synchronizing gestures with speech. Despite recent advancements, existing methods struggle with accurately identifying the…

Computer Vision and Pattern Recognition · Computer Science 2025-09-23 Pinxin Liu , Pengfei Zhang , Hyeongwoo Kim , Pablo Garrido , Ari Shapiro , Kyle Olszewski

Generating realistic human motion is essential for many computer vision and graphics applications. The wide variety of human body shapes and sizes greatly impacts how people move. However, most existing motion models ignore these…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Shashank Tripathi , Omid Taheri , Christoph Lassner , Michael J. Black , Daniel Holden , Carsten Stoll

In this paper, we introduce a simple and novel framework for one-shot audio-driven talking head generation. Unlike prior works that require additional driving sources for controlled synthesis in a deterministic manner, we instead…

Graphics · Computer Science 2022-12-09 Zhentao Yu , Zixin Yin , Deyu Zhou , Duomin Wang , Finn Wong , Baoyuan Wang

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Jiadong Liang , Feng Lu

We propose a new framework for gesture generation, aiming to allow data-driven approaches to produce more semantically rich gestures. Our approach first predicts whether to gesture, followed by a prediction of the gesture properties. Those…

Human-Computer Interaction · Computer Science 2021-09-28 Taras Kucherenko , Rajmund Nagy , Patrik Jonell , Michael Neff , Hedvig Kjellström , Gustav Eje Henter

Gestures are non-verbal but important behaviors accompanying people's speech. While previous methods are able to generate speech rhythm-synchronized gestures, the semantic context of the speech is generally lacking in the gesticulations.…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Yihao Zhi , Xiaodong Cun , Xuelin Chen , Xi Shen , Wen Guo , Shaoli Huang , Shenghua Gao
‹ Prev 1 4 5 6 7 8 10 Next ›