English
Related papers

Related papers: Audio2Gestures: Generating Diverse Gestures from A…

200 papers

State-of-the-art text-to-motion generation models rely on the kinematic-aware, local-relative motion representation popularized by HumanML3D, which encodes motion relative to the pelvis and to the previous frame with built-in redundancy.…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Zichong Meng , Zeyu Han , Xiaogang Peng , Yiming Xie , Huaizu Jiang

We review human evaluation practices in automatic, speech-driven 3D gesture generation and find a lack of standardisation and frequent use of flawed experimental setups. This leads to a situation where it is impossible to know how different…

Human motion synthesis is an important task in computer graphics and computer vision. While focusing on various conditioning signals such as text, action class, or audio to guide the generation process, most existing methods utilize…

Computer Vision and Pattern Recognition · Computer Science 2024-05-14 Kebing Xue , Hyewon Seo

Gesture-driven music generation is an emerging human-computer interaction paradigm for touch-free and expressive musical interaction. However, many existing approaches treat the task as isolated gesture classification or map gestures to…

Multimedia · Computer Science 2026-04-29 Rathinaraja Jeyaraj , Barathi Subramanian , Kapilya Gangadharan , Anand Paul

While the field of co-speech gesture generation has seen significant advances, producing holistic, semantically grounded gestures remains a challenge. Existing approaches rely on external semantic retrieval methods, which limit their…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Lanmiao Liu , Esam Ghaleb , Aslı Özyürek , Zerrin Yumak

How does audio describe the world around us? In this paper, we propose a method for generating an image of a scene from sound. Our method addresses the challenges of dealing with the large gaps that often exist between sight and sound. We…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Kim Sung-Bin , Arda Senocak , Hyunwoo Ha , Andrew Owens , Tae-Hyun Oh

This paper aims to deal with the ignored real-world complexities in prior work on human motion forecasting, emphasizing the social properties of multi-person motion, the diversity of motion and social interactions, and the complexity of…

Computer Vision and Pattern Recognition · Computer Science 2023-06-09 Sirui Xu , Yu-Xiong Wang , Liang-Yan Gui

Synthesizing human motion through learning techniques is becoming an increasingly popular approach to alleviating the requirement of new data capture to produce animations. Learning to move naturally from music, i.e., to dance, is one of…

While previous approaches to 3D human motion generation have achieved notable success, they often rely on extensive training and are limited to specific tasks. To address these challenges, we introduce Motion-Agent, an efficient…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Qi Wu , Yubo Zhao , Yifan Wang , Xinhang Liu , Yu-Wing Tai , Chi-Keung Tang

In this work, we introduce a challenging task for simultaneously generating 3D holistic body motions and singing vocals directly from textual lyrics inputs, advancing beyond existing works that typically address these two modalities in…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jiaben Chen , Xin Yan , Yihang Chen , Siyuan Cen , Zixin Wang , Qinwei Ma , Haoyu Zhen , Kaizhi Qian , Lie Lu , Chuang Gan

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is…

Sound · Computer Science 2025-01-10 Darius Petermann , Mahdi M. Kalayeh

Recent advances in text-to-motion generation using diffusion and autoregressive models have shown promising results. However, these models often suffer from a trade-off between real-time performance, high fidelity, and motion editability.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-29 Ekkasit Pinyoanuntapong , Pu Wang , Minwoo Lee , Chen Chen

Audio-driven portrait animation aims to synthesize portrait videos that are conditioned by given audio. Animating high-fidelity and multimodal video portraits has a variety of applications. Previous methods have attempted to capture…

Computer Vision and Pattern Recognition · Computer Science 2023-07-20 Yunfei Liu , Lijian Lin , Fei Yu , Changyin Zhou , Yu Li

Environmental sounds like footsteps, keyboard typing, or dog barking carry rich information and emotional context, making them valuable for designing haptics in user applications. Existing audio-to-vibration methods, however, rely on…

Human-Computer Interaction · Computer Science 2026-01-27 Yinan Li , Hasti Seifi

We present a learning-based approach for generating binaural audio from mono audio using multi-task learning. Our formulation leverages additional information from two related tasks: the binaural audio generation task and the flipped audio…

Sound · Computer Science 2021-09-03 Sijia Li , Shiguang Liu , Dinesh Manocha

Text to Motion aims to generate human motions from texts. Existing settings rely on limited Action Texts that include action labels, which limits flexibility and practicability in scenarios difficult to describe directly. This paper extends…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Runqi Wang , Caoyuan Ma , Guopeng Li , Hanrui Xu , Yuke Li , Zheng Wang

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Shuling Zhao , Fa-Ting Hong , Xiaoshui Huang , Dan Xu

Recent increase of remote-work, online meeting and tele-operation task makes people find that gesture for avatars and communication robots is more important than we have thought. It is one of the key factors to achieve smooth and natural…

Human-Computer Interaction · Computer Science 2023-09-29 Hitoshi Teshima , Naoki Wake , Diego Thomas , Yuta Nakashima , Hiroshi Kawasaki , Katsushi Ikeuchi

Text-to-motion generation is driven by learning motion representations for semantic alignment with language. Existing methods rely on either continuous or discrete motion representations. However, continuous representations entangle…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Dawei Guan , Di Yang , Chengjie Jin , Jiangtao Wang

Co-speech gesture video generation aims to synthesize realistic, audio-aligned videos of speakers, complete with synchronized facial expressions and body gestures. This task presents challenges due to the significant one-to-many mapping…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Xu Yang , Shaoli Huang , Shenbo Xie , Xuelin Chen , Yifei Liu , Changxing Ding