English
Related papers

Related papers: C2G2: Controllable Co-speech Gesture Generation wi…

200 papers

We propose a zero-shot approach for generating consistent videos of animated characters based on Text-to-Image (T2I) diffusion models. Existing Text-to-Video (T2V) methods are expensive to train and require large-scale video datasets to…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Abdelrahman Eldesokey , Peter Wonka

The dyadic reaction generation task involves synthesizing responsive facial reactions that align closely with the behaviors of a conversational partner, enhancing the naturalness and effectiveness of human-like interaction simulations. This…

Machine Learning · Computer Science 2025-05-14 Minh-Duc Nguyen , Hyung-Jeong Yang , Soo-Hyung Kim , Ji-Eun Shin , Seung-Won Kim

This paper presents a novel framework for speech-driven gesture production, applicable to virtual agents to enhance human-computer interaction. Specifically, we extend recent deep-learning-based, data-driven methods for speech-driven…

Computer Vision and Pattern Recognition · Computer Science 2021-04-07 Taras Kucherenko , Dai Hasegawa , Naoshi Kaneko , Gustav Eje Henter , Hedvig Kjellström

We propose LiveGesture, the first fully streamable, speech-driven full-body gesture generation framework that operates with zero look-ahead and supports arbitrary sequence length. Unlike existing co-speech gesture methods, which are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Muhammad Usama Saleem , Mayur Jagdishbhai Patel , Ekkasit Pinyoanuntapong , Zhongxing Qin , Li Yang , Hongfei Xue , Ahmed Helmy , Chen Chen , Pu Wang

Human communication seamlessly integrates speech and bodily motion, where hand gestures naturally complement vocal prosody to express intent, emotion, and emphasis. While recent text-to-speech (TTS) systems have begun incorporating…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-23 Lokesh Kumar , Nirmesh Shah , Ashishkumar P. Gudmalwar , Pankaj Wasnik

Current audio-driven 3D head generation methods mainly focus on single-speaker scenarios, lacking natural, bidirectional listen-and-speak interaction. Achieving seamless conversational behavior, where speaking and listening states…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Lei Zhu , Lijian Lin , Ye Zhu , Jiahao Wu , Xuehan Hou , Yu Li , Yunfei Liu , Jie Chen

With read-aloud speech synthesis achieving high naturalness scores, there is a growing research interest in synthesising spontaneous speech. However, human spontaneous face-to-face conversation has both spoken and non-verbal aspects (here,…

Audio and Speech Processing · Electrical Eng. & Systems 2023-09-15 Shivam Mehta , Siyang Wang , Simon Alexanderson , Jonas Beskow , Éva Székely , Gustav Eje Henter

Embodied conversational agents (ECA) are often designed to produce nonverbal behavior to complement or enhance their verbal communication. One such form of nonverbal behavior is co-speech gesturing, which involves movements that the agent…

Human-Computer Interaction · Computer Science 2022-03-02 Pieter Wolfert , Nicole Robinson , Tony Belpaeme

People may perform diverse gestures affected by various mental and physical factors when speaking the same sentences. This inherent one-to-many relationship makes co-speech gesture generation from audio particularly challenging.…

Computer Vision and Pattern Recognition · Computer Science 2023-01-18 Jing Li , Di Kang , Wenjie Pei , Xuefei Zhe , Ying Zhang , Linchao Bao , Zhenyu He

Speech-driven talking heads have recently emerged and enable interactive avatars. However, real-world applications are limited, as current methods achieve high visual fidelity but slow or fast yet temporally unstable. Diffusion methods…

Computer Vision and Pattern Recognition · Computer Science 2025-12-12 Madhav Agarwal , Mingtian Zhang , Laura Sevilla-Lara , Steven McDonagh

Audio-driven talking head generation is critical for applications such as virtual assistants, video games, and films, where natural lip movements are essential. Despite progress in this field, challenges remain in producing both consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Yucheng Wang , Dan Xu

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing methods ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Chao Xu , Shaoting Zhu , Junwei Zhu , Tianxin Huang , Jiangning Zhang , Ying Tai , Yong Liu

Speech-driven 3D talking face method should offer both accurate lip synchronization and controllable expressions. Previous methods solely adopt discrete emotion labels to globally control expressions throughout sequences while limiting…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Hejia Chen , Haoxian Zhang , Shoulong Zhang , Xiaoqiang Liu , Sisi Zhuang , Yuan Zhang , Pengfei Wan , Di Zhang , Shuai Li

Text-to-image generation tasks have driven remarkable advances in diverse media applications, yet most focus on single-turn scenarios and struggle with iterative, multi-turn creative tasks. Recent dialogue-based systems attempt to bridge…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Shichao Ma , Yunhe Guo , Jiahao Su , Qihe Huang , Zhengyang Zhou , Yang Wang

In real-life conversations, the content is diverse, and there exists the one-to-many problem that requires diverse generation. Previous studies attempted to introduce discrete or Gaussian-based continuous latent variables to address the…

Computation and Language · Computer Science 2024-04-11 Jianxiang Xiang , Zhenhua Liu , Haodong Liu , Yin Bai , Jia Cheng , Wenliang Chen

Despite the remarkable process of talking-head-based avatar-creating solutions, directly generating anchor-style videos with full-body motions remains challenging. In this study, we propose Make-Your-Anchor, a novel system necessitating…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Ziyao Huang , Fan Tang , Yong Zhang , Xiaodong Cun , Juan Cao , Jintao Li , Tong-Yee Lee

Recent advances in audio-driven talking head generation have achieved impressive results in lip synchronization and emotional expression. However, they largely overlook the crucial task of facial attribute editing. This capability is…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Guanwen Feng , Zhiyuan Ma , Yunan Li , Jiahao Yang , Junwei Jing , Qiguang Miao

Real-time, streaming interactive avatars represent a critical yet challenging goal in digital human research. Although diffusion-based human avatar generation methods achieve remarkable success, their non-causal architecture and high…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Zhiyao Sun , Ziqiao Peng , Yifeng Ma , Yi Chen , Zhengguang Zhou , Zixiang Zhou , Guozhen Zhang , Youliang Zhang , Yuan Zhou , Qinglin Lu , Yong-Jin Liu

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural motion for an…

Computer Vision and Pattern Recognition · Computer Science 2024-01-04 Evonne Ng , Javier Romero , Timur Bagautdinov , Shaojie Bai , Trevor Darrell , Angjoo Kanazawa , Alexander Richard

The effective communication of procedural knowledge remains a significant challenge in natural language processing (NLP), as purely textual instructions often fail to convey complex physical actions and spatial relationships. We address…

Computation and Language · Computer Science 2025-05-23 Jing Bi , Pinxin Liu , Ali Vosoughi , Jiarui Wu , Jinxi He , Chenliang Xu