English
Related papers

Related papers: LiveGesture Streamable Co-Speech Gesture Generatio…

200 papers

The generation of temporally consistent, high-fidelity driving videos over extended horizons presents a fundamental challenge in autonomous driving world modeling. Existing approaches often suffer from error accumulation and feature…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jiamin Wang , Yichen Yao , Xiang Feng , Hang Wu , Yaming Wang , Qingqiu Huang , Yuexin Ma , Xinge Zhu

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

Existing humanoid control systems often rely on teleoperation or modular generation pipelines that separate language understanding from physical execution. However, the former is entirely human-driven, and the latter lacks tight alignment…

Robotics · Computer Science 2025-11-25 Yuxuan Wang , Haobin Jiang , Shiqing Yao , Ziluo Ding , Zongqing Lu

We present visual action prompts, a unified action representation for action-to-video generation of complex high-DoF interactions while maintaining transferable visual dynamics across domains. Action-driven video generation faces a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Yuang Wang , Chao Wen , Haoyu Guo , Sida Peng , Minghan Qin , Hujun Bao , Xiaowei Zhou , Ruizhen Hu

This paper introduces Stereo-Talker, a novel one-shot audio-driven human video synthesis system that generates 3D talking videos with precise lip synchronization, expressive body gestures, temporally consistent photo-realistic quality, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-02 Xiang Deng , Youxin Pang , Xiaochen Zhao , Chao Xu , Lizhen Wang , Hongjiang Xiao , Shi Yan , Hongwen Zhang , Yebin Liu

Modeling virtual agents with behavior style is one factor for personalizing human agent interaction. We propose an efficient yet effective machine learning approach to synthesize gestures driven by prosodic features and text in the style of…

Sound · Computer Science 2022-08-04 Mireille Fares , Michele Grimaldi , Catherine Pelachaud , Nicolas Obin

Audio-driven human gesture synthesis is a crucial task with broad applications in virtual avatars, human-computer interaction, and creative content generation. Despite notable progress, existing methods often produce gestures that are…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Xukun Zhou , Fengxin Li , Ming Chen , Yan Zhou , Pengfei Wan , Di Zhang , Yeying Jin , Zhaoxin Fan , Hongyan Liu , Jun He

Text-driven motion generation offers a powerful and intuitive way to create human movements directly from natural language. By removing the need for predefined motion inputs, it provides a flexible and accessible approach to controlling…

Computer Vision and Pattern Recognition · Computer Science 2025-05-15 Ali Rida Sahili , Najett Neji , Hedi Tabia

Diffusion-based video generation technology has advanced significantly, catalyzing a proliferation of research in human animation. However, the majority of these studies are confined to same-modality driving settings, with cross-modality…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Gaojie Lin , Jianwen Jiang , Chao Liang , Tianyun Zhong , Jiaqi Yang , Yanbo Zheng

Large Language Models have shown remarkable efficacy in generating streaming data such as text and audio, thanks to their temporally uni-directional attention mechanism, which models correlations between the current token and previous…

Computer Vision and Pattern Recognition · Computer Science 2024-07-12 Zhening Xing , Gereon Fox , Yanhong Zeng , Xingang Pan , Mohamed Elgharib , Christian Theobalt , Kai Chen

Co-speech gestures are crucial non-verbal cues that enhance speech clarity and expressiveness in human communication, which have attracted increasing attention in multimodal research. While the existing methods have made strides in gesture…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Hongye Cheng , Tianyu Wang , Guangsi Shi , Zexing Zhao , Yanwei Fu

In this paper, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strategy sets up…

Computer Vision and Pattern Recognition · Computer Science 2024-01-09 Xuehao Gao , Yang Yang , Zhenyu Xie , Shaoyi Du , Zhongqian Sun , Yang Wu

Autoregressive models for video generation typically operate frame-by-frame, extending next-token prediction from language to video's temporal dimension. We question that unlike word as token is universally agreed in language if frame is a…

Computer Vision and Pattern Recognition · Computer Science 2025-09-30 Sucheng Ren , Chen Chen , Zhenbang Wang , Liangchen Song , Xiangxin Zhu , Alan Yuille , Yinfei Yang , Jiasen Lu

How can we teach robots or virtual assistants to gesture naturally? Can we go further and adapt the gesturing style to follow a specific speaker? Gestures that are naturally timed with corresponding speech during human communication are…

Computer Vision and Pattern Recognition · Computer Science 2020-07-27 Chaitanya Ahuja , Dong Won Lee , Yukiko I. Nakano , Louis-Philippe Morency

This paper focuses on the task of speech-driven 3D facial animation, which aims to generate realistic and synchronized facial motions driven by speech inputs. Recent methods have employed audio-conditioned diffusion models for 3D facial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Yifan Yang , Zhi Cen , Sida Peng , Xiangwei Chen , Yifu Deng , Xinyu Zhu , Fan Jia , Xiaowei Zhou , Hujun Bao

Portrait Animation aims to synthesize a lifelike video from a single source image, using it as an appearance reference, with motion (i.e., facial expressions and head pose) derived from a driving video, audio, text, or generation. Instead…

Computer Vision and Pattern Recognition · Computer Science 2025-03-03 Jianzhu Guo , Dingyun Zhang , Xiaoqiang Liu , Zhizhou Zhong , Yuan Zhang , Pengfei Wan , Di Zhang

Modeling and synthesizing complex hand-object interactions remains a significant challenge, even for state-of-the-art physics engines. Conventional simulation-based approaches rely on explicitly defined rigid object models and pre-scripted…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Zisu Li , Hengye Lyu , Jiaxin Shi , Yufeng Zeng , Mingming Fan , Hanwang Zhang , Chen Liang

Achieving a balance between high-fidelity visual quality and low-latency streaming remains a formidable challenge in audio-driven portrait generation. Existing large-scale models often suffer from prohibitive computational costs, while…

Computer Vision and Pattern Recognition · Computer Science 2026-02-12 Tan Yu , Qian Qiao , Le Shen , Ke Zhou , Jincheng Hu , Dian Sheng , Bo Hu , Haoming Qin , Jun Gao , Changhai Zhou , Shunshun Yin , Siyuan Liu

Talking face generation aims to synthesize a face video with precise lip synchronization as well as a smooth transition of facial motion over the entire video via the given speech clip and facial image. Most existing methods mainly focus on…

Computer Vision and Pattern Recognition · Computer Science 2020-05-14 Hao Zhu , Huaibo Huang , Yi Li , Aihua Zheng , Ran He

Autoregressive (AR) diffusion models offer a promising framework for sequential generation tasks such as video synthesis by combining diffusion modeling with causal inference. Although they support streaming generation, existing AR…

Computer Vision and Pattern Recognition · Computer Science 2026-03-20 Dingcheng Zhen , Xu Zheng , Ruixin Zhang , Zhiqi Jiang , Yichao Yan , Ming Tao , Shunshun Yin