中文
相关论文

相关论文: Diffusion-Based Co-Speech Gesture Generation Using…

200 篇论文

In this paper, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strategy sets up…

计算机视觉与模式识别 · 计算机科学 2024-01-09 Xuehao Gao , Yang Yang , Zhenyu Xie , Shaoyi Du , Zhongqian Sun , Yang Wu

Synthesizing realistic co-speech gestures is an important and yet unsolved problem for creating believable motions that can drive a humanoid robot to interact and communicate with human users. Such capability will improve the impressions of…

计算机视觉与模式识别 · 计算机科学 2023-03-24 Shuhong Lu , Youngwoo Yoon , Andrew Feng

Embodied human communication encompasses both verbal (speech) and non-verbal information (e.g., gesture and head movements). Recent advances in machine learning have substantially improved the technologies for generating synthetic versions…

机器学习 · 计算机科学 2021-01-15 Simon Alexanderson , Éva Székely , Gustav Eje Henter , Taras Kucherenko , Jonas Beskow

This paper presents a novel framework for automatic speech-driven gesture generation, applicable to human-agent interaction including both virtual agents and robots. Specifically, we extend recent deep-learning-based, data-driven methods…

人机交互 · 计算机科学 2019-06-12 Taras Kucherenko , Dai Hasegawa , Gustav Eje Henter , Naoshi Kaneko , Hedvig Kjellström

Current evaluation practices in speech-driven gesture generation lack standardisation and focus on aspects that are easy to measure over aspects that actually matter. This leads to a situation where it is impossible to know what is the…

Human communication is multimodal, with speech and gestures tightly coupled, yet most computational methods for generating speech and gestures synthesize them sequentially, weakening synchrony and prosody alignment. We introduce Gelina, a…

Co-speech gesture generation has significantly advanced human-computer interaction, yet speaker movements remain constrained due to the omission of text-driven non-spontaneous gestures (e.g., bowing while talking). Existing methods face two…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Fengyi Fang , Sicheng Yang , Wenming Yang

The automatic generation of stylized co-speech gestures has recently received increasing attention. Previous systems typically allow style control via predefined text labels or example motion clips, which are often not flexible enough to…

计算机视觉与模式识别 · 计算机科学 2023-10-17 Tenglong Ao , Zeyi Zhang , Libin Liu

Conversational scenarios are very common in real-world settings, yet existing co-speech motion synthesis approaches often fall short in these contexts, where one person's audio and gestures will influence the other's responses.…

声音 · 计算机科学 2024-12-04 Mingyi Shi , Dafei Qin , Leo Ho , Zhouyingcheng Liao , Yinghao Huang , Junichi Yamagishi , Taku Komura

Co-speech gestures are crucial non-verbal cues that enhance speech clarity and expressiveness in human communication, which have attracted increasing attention in multimodal research. While the existing methods have made strides in gesture…

计算机视觉与模式识别 · 计算机科学 2025-03-04 Hongye Cheng , Tianyu Wang , Guangsi Shi , Zexing Zhao , Yanwei Fu

Generative AI has demonstrated impressive performance in various fields, among which speech synthesis is an interesting direction. With the diffusion model as the most popular generative model, numerous works have attempted two active…

We propose a simple and novel method for generating 3D human motion from complex natural language sentences, which describe different velocity, direction and composition of all kinds of actions. Different from existing methods that use…

计算机视觉与模式识别 · 计算机科学 2023-04-17 Zhiyuan Ren , Zhihong Pan , Xin Zhou , Le Kang

Gestures that accompany speech are an essential part of natural and efficient embodied human communication. The automatic generation of such co-speech gestures is a long-standing problem in computer animation and is considered an enabling…

图形学 · 计算机科学 2023-04-11 Simbarashe Nyatsanga , Taras Kucherenko , Chaitanya Ahuja , Gustav Eje Henter , Michael Neff

We propose DiffSep, a new single channel source separation method based on score-matching of a stochastic differential equation (SDE). We craft a tailored continuous time diffusion-mixing process starting from the separated sources and…

音频与语音处理 · 电气工程与系统科学 2022-11-03 Robin Scheibler , Youna Ji , Soo-Whan Chung , Jaeuk Byun , Soyeon Choe , Min-Seok Choi

Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate…

计算机视觉与模式识别 · 计算机科学 2025-04-07 M. Hamza Mughal , Rishabh Dabral , Merel C. J. Scholman , Vera Demberg , Christian Theobalt

While the field of co-speech gesture generation has seen significant advances, producing holistic, semantically grounded gestures remains a challenge. Existing approaches rely on external semantic retrieval methods, which limit their…

计算机视觉与模式识别 · 计算机科学 2026-03-30 Lanmiao Liu , Esam Ghaleb , Aslı Özyürek , Zerrin Yumak

As text-to-speech technologies achieve remarkable naturalness in read-aloud tasks, there is growing interest in multimodal synthesis of verbal and non-verbal communicative behaviour, such as spontaneous speech and associated body gestures.…

音频与语音处理 · 电气工程与系统科学 2024-01-11 Shivam Mehta , Ruibo Tu , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter

Animating virtual characters with holistic co-speech gestures is a challenging but critical task. Previous systems have primarily focused on the weak correlation between audio and gestures, leading to physically unnatural outcomes that…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Yongkang Cheng , Shaoli Huang

Human communication combines speech with expressive nonverbal cues such as hand gestures that serve manifold communicative functions. Yet, current generative gesture generation approaches are restricted to simple, repetitive beat gestures…

人机交互 · 计算机科学 2025-10-21 Hendric Voss , Stefan Kopp

Audio-driven talking video generation has advanced significantly, but existing methods often depend on video-to-video translation techniques and traditional generative networks like GANs and they typically generate taking heads and…

计算机视觉与模式识别 · 计算机科学 2024-09-13 Steven Hogue , Chenxu Zhang , Hamza Daruger , Yapeng Tian , Xiaohu Guo