English
Related papers

Related papers: Grounded Gesture Generation: Language, Motion, and…

200 papers

Recent progress in large models has led to significant advances in unified multimodal generation and understanding. However, the development of models that unify motion-language generation and understanding remains largely underexplored.…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Zekun Li , Sizhe An , Chengcheng Tang , Chuan Guo , Ivan Shugurov , Linguang Zhang , Amy Zhao , Srinath Sridhar , Lingling Tao , Abhay Mittal

Video generation has witnessed great success recently, but their application in generating long videos still remains challenging due to the difficulty in maintaining the temporal consistency of generated videos and the high memory cost…

Computer Vision and Pattern Recognition · Computer Science 2024-07-19 Wei Feng , Xin Wang , Hong Chen , Zeyang Zhang , Wenwu Zhu

Language-driven dexterous grasp generation requires the models to understand task semantics, 3D geometry, and complex hand-object interactions. While vision-language models have been applied to this problem, existing approaches directly map…

Robotics · Computer Science 2026-04-28 Junha Lee , Eunha Park , Minsu Cho

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

3D vision-language grounding, which focuses on aligning language with the 3D physical environment, stands as a cornerstone in the development of embodied agents. In comparison to recent advancements in the 2D domain, grounding language in…

Computer Vision and Pattern Recognition · Computer Science 2024-09-25 Baoxiong Jia , Yixin Chen , Huangyue Yu , Yan Wang , Xuesong Niu , Tengyu Liu , Qing Li , Siyuan Huang

The generation of realistic and contextually relevant co-speech gestures is a challenging yet increasingly important task in the creation of multimodal artificial agents. Prior methods focused on learning a direct correspondence between…

Human-Computer Interaction · Computer Science 2023-05-09 Hendric Voß , Stefan Kopp

Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-23 Lokesh Kumar , Nirmesh Shah , Ashishkumar P. Gudmalwar , Pankaj Wasnik

In this work, we focus on the problem of grounding language by training an agent to follow a set of natural language instructions and navigate to a target object in an environment. The agent receives visual information through raw pixels…

Computation and Language · Computer Science 2018-12-27 Akilesh B , Abhishek Sinha , Mausoom Sarkar , Balaji Krishnamurthy

In this work, we present Semantic Gesticulator, a novel framework designed to synthesize realistic gestures accompanying speech with strong semantic correspondence. Semantically meaningful gestures are crucial for effective non-verbal…

Graphics · Computer Science 2025-10-23 Zeyi Zhang , Tenglong Ao , Yuyao Zhang , Qingzhe Gao , Chuan Lin , Baoquan Chen , Libin Liu

This work aims to generate natural and diverse group motions of multiple humans from textual descriptions. While single-person text-to-motion generation is extensively studied, it remains challenging to synthesize motions for more than one…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Mengyi Shan , Lu Dong , Yutao Han , Yuan Yao , Tao Liu , Ifeoma Nwogu , Guo-Jun Qi , Mitch Hill

Automatic gesture synthesis from speech is a topic that has attracted researchers for applications in remote communication, video games and Metaverse. Learning the mapping between speech and 3D full-body gestures is difficult due to the…

Computer Vision and Pattern Recognition · Computer Science 2023-10-12 Kunkun Pang , Dafei Qin , Yingruo Fan , Julian Habekost , Takaaki Shiratori , Junichi Yamagishi , Taku Komura

We introduce a technique for multi-document grounded multi-turn synthetic dialog generation that incorporates three main ideas. First, we control the overall dialog flow using taxonomy-driven user queries that are generated with…

Computation and Language · Computer Science 2024-09-19 Young-Suk Lee , Chulaka Gunasekara , Danish Contractor , Ramón Fernandez Astudillo , Radu Florian

Common grounding is the process of creating and maintaining mutual understandings, which is a critical aspect of sophisticated human communication. While various task settings have been proposed in existing literature, they mostly focus on…

Computation and Language · Computer Science 2021-06-01 Takuma Udagawa , Akiko Aizawa

Recent advancements in multimodal Human-Robot Interaction (HRI) datasets have highlighted the fusion of speech and gesture, expanding robots' capabilities to absorb explicit and implicit HRI insights. However, existing speech-gesture HRI…

Robotics · Computer Science 2024-03-05 Snehesh Shrestha , Yantian Zha , Saketh Banagiri , Ge Gao , Yiannis Aloimonos , Cornelia Fermuller

Denoising diffusion models have shown great promise in human motion synthesis conditioned on natural language descriptions. However, integrating spatial constraints, such as pre-defined motion trajectories and obstacles, remains a challenge…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Korrawe Karunratanakul , Konpat Preechakul , Supasorn Suwajanakorn , Siyu Tang

We introduce the concept of "empathic grounding" in conversational agents as an extension of Clark's conceptualization of grounding in conversation in which the grounding criterion includes listener empathy for the speaker's affective…

Human-Computer Interaction · Computer Science 2024-07-03 Mehdi Arjmand , Farnaz Nouraei , Ian Steenstra , Timothy Bickmore

Generating vivid and emotional 3D co-speech gestures is crucial for virtual avatar animation in human-machine interaction applications. While the existing methods enable generating the gestures to follow a single emotion label, they…

Computer Vision and Pattern Recognition · Computer Science 2024-03-28 Xingqun Qi , Jiahao Pan , Peng Li , Ruibin Yuan , Xiaowei Chi , Mengfei Li , Wenhan Luo , Wei Xue , Shanghang Zhang , Qifeng Liu , Yike Guo

Speech-driven gesture synthesis is a field of growing interest in virtual human creation. However, a critical challenge is the inherent intricate one-to-many mapping between speech and gestures. Previous studies have explored and achieved…

Graphics · Computer Science 2023-02-03 Fan Zhang , Naye Ji , Fuxing Gao , Yongping Li

Text-to-motion generation has experienced remarkable progress in recent years. However, current approaches remain limited to synthesizing motion from short or general text prompts, primarily due to dataset constraints. This limitation…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Chuan Guo , Inwoo Hwang , Jian Wang , Bing Zhou

Large language models (LLMs) have unified diverse linguistic tasks within a single framework, yet such unification remains unexplored in human motion generation. Existing methods are confined to isolated tasks, limiting flexibility for…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Wendong Bu , Kaihang Pan , Yuze Lin , Jiacheng Li , Kai Shen , Wenqiao Zhang , Juncheng Li , Jun Xiao , Siliang Tang