中文
相关论文

相关论文: Contextual Gesture: Co-Speech Gesture Video Genera…

200 篇论文

Large Language Model (LLM)-driven digital humans have sparked a series of recent studies on co-speech gesture generation systems. However, existing approaches struggle with real-time synthesis and long-text comprehension. This paper…

图形学 · 计算机科学 2025-06-03 Yueqian Guo , Tianzhao Li , Xin Lyu , Jiehaolin Chen , Zhaohan Wang , Sirui Xiao , Yurun Chen , Yezi He , Helin Li , Fan Zhang

Pointing is a key mode of interaction with robots, yet most prior work has focused on recognition rather than generation. We present a motion capture dataset of human pointing gestures covering diverse styles, handedness, and spatial…

机器人学 · 计算机科学 2025-09-17 Anna Deichler , Siyang Wang , Simon Alexanderson , Jonas Beskow

In this paper, we propose a novel approach to convert given speech audio to a photo-realistic speaking video of a specific person, where the output video has synchronized, realistic, and expressive rich body dynamics. We achieve this by…

计算机视觉与模式识别 · 计算机科学 2020-10-12 Miao Liao , Sibo Zhang , Peng Wang , Hao Zhu , Xinxin Zuo , Ruigang Yang

Co-speech gestures, gestures that accompany speech, play an important role in human communication. Automatic co-speech gesture generation is thus a key enabling technology for embodied conversational agents (ECAs), since humans expect ECAs…

人机交互 · 计算机科学 2021-02-24 Taras Kucherenko , Patrik Jonell , Youngwoo Yoon , Pieter Wolfert , Gustav Eje Henter

Non-verbal communication often comprises of semantically rich gestures that help convey the meaning of an utterance. Producing such semantic co-speech gestures has been a major challenge for the existing neural systems that can generate…

计算机视觉与模式识别 · 计算机科学 2025-04-07 M. Hamza Mughal , Rishabh Dabral , Merel C. J. Scholman , Vera Demberg , Christian Theobalt

This paper proposes a network architecture to perform variable length semantic video generation using captions. We adopt a new perspective towards video generation where we allow the captions to be combined with the long-term and short-term…

计算机视觉与模式识别 · 计算机科学 2017-11-17 Tanya Marwah , Gaurav Mittal , Vineeth N. Balasubramanian

Video generation has achieved remarkable progress with the introduction of diffusion models, which have significantly improved the quality of generated videos. However, recent research has primarily focused on scaling up model training,…

计算机视觉与模式识别 · 计算机科学 2025-01-16 Chenyang Si , Weichen Fan , Zhengyao Lv , Ziqi Huang , Yu Qiao , Ziwei Liu

Recent increase of remote-work, online meeting and tele-operation task makes people find that gesture for avatars and communication robots is more important than we have thought. It is one of the key factors to achieve smooth and natural…

人机交互 · 计算机科学 2023-09-29 Hitoshi Teshima , Naoki Wake , Diego Thomas , Yuta Nakashima , Hiroshi Kawasaki , Katsushi Ikeuchi

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Lele Chen , Guofeng Cui , Celong Liu , Zhong Li , Ziyi Kou , Yi Xu , Chenliang Xu

Text-to-speech and co-speech gesture synthesis have until now been treated as separate areas by two different research communities, and applications merely stack the two technologies using a simple system-level pipeline. This can lead to…

人机交互 · 计算机科学 2021-08-27 Siyang Wang , Simon Alexanderson , Joakim Gustafson , Jonas Beskow , Gustav Eje Henter , Éva Székely

This paper reports on the second GENEA Challenge to benchmark data-driven automatic co-speech gesture generation. Participating teams used the same speech and motion dataset to build gesture-generation systems. Motion generated by all these…

Co-speech gestures enhance interaction experiences between humans as well as between humans and robots. Existing robots use rule-based speech-gesture association, but this requires human labor and prior knowledge of experts to be…

机器人学 · 计算机科学 2018-10-31 Youngwoo Yoon , Woo-Ri Ko , Minsu Jang , Jaeyeon Lee , Jaehong Kim , Geehyuk Lee

Creating a diverse and comprehensive dataset of hand gestures for dynamic human-machine interfaces in the automotive domain can be challenging and time-consuming. To overcome this challenge, we propose using synthetic gesture datasets…

计算机视觉与模式识别 · 计算机科学 2024-08-05 Amr Gomaa , Robin Zitt , Guillermo Reyes , Antonio Krüger

The objective of this paper is to jointly synthesize interactive videos and conversational speech from text and reference images. With the ultimate goal of building human-like conversational systems, recent studies have explored talking or…

计算机视觉与模式识别 · 计算机科学 2025-12-24 Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Joon Son Chung , Shinji Watanabe

Recent advancements in zero-shot text-to-speech (TTS) modeling have led to significant strides in generating high-fidelity and diverse speech. However, dialogue generation, along with achieving human-like naturalness in speech, continues to…

音频与语音处理 · 电气工程与系统科学 2024-12-17 Leying Zhang , Yao Qian , Long Zhou , Shujie Liu , Dongmei Wang , Xiaofei Wang , Midia Yousefi , Yanmin Qian , Jinyu Li , Lei He , Sheng Zhao , Michael Zeng

We devise a cascade GAN approach to generate talking face video, which is robust to different face shapes, view angles, facial characteristics, and noisy audio conditions. Instead of learning a direct mapping from audio to video frames, we…

计算机视觉与模式识别 · 计算机科学 2019-05-13 Lele Chen , Ross K. Maddox , Zhiyao Duan , Chenliang Xu

This study mainly explores the application of natural gesture recognition based on computer vision in human-computer interaction, aiming to improve the fluency and naturalness of human-computer interaction through gesture recognition…

计算机视觉与模式识别 · 计算机科学 2024-12-25 Fenghua Shao , Tong Zhang , Shang Gao , Qi Sun , Liuqingqing Yang

Scene-consistent video generation aims to create videos that explore 3D scenes based on a camera trajectory. Previous methods rely on video generation models with external memory for consistency, or iterative 3D reconstruction and…

计算机视觉与模式识别 · 计算机科学 2026-02-26 JiaKui Hu , Jialun Liu , Liying Yang , Xinliang Zhang , Kaiwen Li , Shuang Zeng , Yuanwei Li , Haibin Huang , Chi Zhang , Yanye Lu

We propose DiffSHEG, a Diffusion-based approach for Speech-driven Holistic 3D Expression and Gesture generation with arbitrary length. While previous works focused on co-speech gesture or expression generation individually, the joint…

声音 · 计算机科学 2024-04-09 Junming Chen , Yunfei Liu , Jianan Wang , Ailing Zeng , Yu Li , Qifeng Chen

Recent advances in scene-based video generation enable coherent visual narratives from structured prompts, yet a key aspect of storytelling -- character-driven dialogue and speech -- remains underexplored. We present a modular pipeline that…

计算机视觉与模式识别 · 计算机科学 2026-05-20 Taewon Kang , Ming C. Lin