English
Related papers

Related papers: InterDyad: Interactive Dyadic Speech-to-Video Gene…

200 papers

While previous audio-driven talking head generation (THG) methods generate head poses from driving audio, the generated poses or lips cannot match the audio well or are not editable. In this study, we propose \textbf{PoseTalk}, a THG system…

Computer Vision and Pattern Recognition · Computer Science 2024-09-05 Jun Ling , Yiwen Wang , Han Xue , Rong Xie , Li Song

We present VideoReTalking, a new system to edit the faces of a real-world talking head video according to input audio, producing a high-quality and lip-syncing output video even with a different emotion. Our system disentangles this…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Kun Cheng , Xiaodong Cun , Yong Zhang , Menghan Xia , Fei Yin , Mingrui Zhu , Xuan Wang , Jue Wang , Nannan Wang

Significant advances have been made in human-centric video generation, yet the joint video-depth generation problem remains underexplored. Most existing monocular depth estimation methods may not generalize well to synthesized images or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Yuanhao Zhai , Kevin Lin , Linjie Li , Chung-Ching Lin , Jianfeng Wang , Zhengyuan Yang , David Doermann , Junsong Yuan , Zicheng Liu , Lijuan Wang

In today's globalized world, effective communication with people from diverse linguistic backgrounds has become increasingly crucial. While traditional methods of language translation, such as written text or voice-only translations, can…

Computation and Language · Computer Science 2023-09-21 Prottay Kumar Adhikary , Bandaru Sugandhi , Subhojit Ghimire , Santanu Pal , Partha Pakray

This paper addresses a novel task of anticipating 3D human-object interactions (HOIs). Most existing research on HOI synthesis lacks comprehensive whole-body interactions with dynamic objects, e.g., often limited to manipulating small or…

Computer Vision and Pattern Recognition · Computer Science 2023-09-01 Sirui Xu , Zhengyuan Li , Yu-Xiong Wang , Liang-Yan Gui

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow…

Computer Vision and Pattern Recognition · Computer Science 2025-09-05 Xuangeng Chu , Nabarun Goswami , Ziteng Cui , Hanqin Wang , Tatsuya Harada

Despite numerous completed studies, achieving high fidelity talking face generation with highly synchronized lip movements corresponding to arbitrary audio remains a significant challenge in the field. The shortcomings of published studies…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Juan Zhang , Jiahao Chen , Cheng Wang , Zhiwang Yu , Tangquan Qi , Di Wu

Recent advances in video generation have shown remarkable potential for constructing world simulators. However, current models still struggle to produce physically consistent results, particularly when handling large-scale or complex…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Zijun Wang , Panwen Hu , Jing Wang , Terry Jingchen Zhang , Yuhao Cheng , Long Chen , Yiqiang Yan , Zutao Jiang , Hanhui Li , Xiaodan Liang

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Weizhi Zhong , Chaowei Fang , Yinqi Cai , Pengxu Wei , Gangming Zhao , Liang Lin , Guanbin Li

The generation of emotional talking faces from a single portrait image remains a significant challenge. The simultaneous achievement of expressive emotional talking and accurate lip-sync is particularly difficult, as expressiveness is often…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Chenxu Zhang , Chao Wang , Jianfeng Zhang , Hongyi Xu , Guoxian Song , You Xie , Linjie Luo , Yapeng Tian , Xiaohu Guo , Jiashi Feng

Researchers have shown a growing interest in Audio-driven Talking Head Generation. The primary challenge in talking head generation is achieving audio-visual coherence between the lips and the audio, known as lip synchronization. This paper…

Sound · Computer Science 2026-02-03 Zhipeng Chen , Xinheng Wang , Lun Xie , Haijie Yuan , Hang Pan

Multi-subject image generation requires seamlessly harmonizing multiple reference identities within a coherent scene. However, existing methods relying on rigid spatial masks or localized attention often struggle with the…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Honghao Cai , Xiangyuan Wang , Jing Li , Yunhao Bai , Tianze Zhou , Haohua Chen , Chao Hui , Changhao Qiao , Runqi Wang , Sijie Xu , Yuyang Hao , Zezhou Cui , Yuyuan Yang , Wei Zhu , Yibo Chen , Xu Tang , Yao Hu , Zhen Li

Audio-driven talking-head generation has advanced rapidly with diffusion-based generative models, yet producing temporally coherent videos with fine-grained motion control remains challenging. We propose DEMO, a flow-matching generative…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Peiyin Chen , Zhuowei Yang , Hui Feng , Sheng Jiang , Rui Yan

In this work, we propose TediGAN, a novel framework for multi-modal image generation and manipulation with textual descriptions. The proposed method consists of three components: StyleGAN inversion module, visual-linguistic similarity…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Weihao Xia , Yujiu Yang , Jing-Hao Xue , Baoyuan Wu

Social intelligence, the ability to interpret emotions, intentions, and behaviors, is essential for effective communication and adaptive responses. As robots and AI systems become more prevalent in caregiving, healthcare, and education, the…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Erika Mori , Yue Qiu , Hirokatsu Kataoka , Yoshimitsu Aoki

We present a multimodal learning-based method to simultaneously synthesize co-speech facial expressions and upper-body gestures for digital characters using RGB video data captured using commodity cameras. Our approach learns from sparse…

Computer Vision and Pattern Recognition · Computer Science 2024-11-26 Uttaran Bhattacharya , Aniket Bera , Dinesh Manocha

Bilingual text-to-motion generation, which synthesizes 3D human motions from bilingual text inputs, holds immense potential for cross-linguistic applications in gaming, film, and robotics. However, this task faces critical challenges: the…

Computer Vision and Pattern Recognition · Computer Science 2025-08-04 Wanjiang Weng , Xiaofeng Tan , Hongsong Wang , Pan Zhou

Gesture recognition research, unlike NLP, continues to face acute data scarcity, with progress constrained by the need for costly human recordings or image processing approaches that cannot generate authentic variability in the gestures…

Computer Vision and Pattern Recognition · Computer Science 2026-04-17 Hassan Ali , Doreen Jirak , Luca Müller , Stefan Wermter

Speech-driven 3D facial animation is challenging due to the scarcity of large-scale visual-audio datasets despite extensive research. Most prior works, typically focused on learning regression models on a small dataset using the method of…

Computer Vision and Pattern Recognition · Computer Science 2024-01-26 Inkyu Park , Jaewoong Cho

In recent years, audio-driven 3D facial animation has gained significant attention, particularly in applications such as virtual reality, gaming, and video conferencing. However, accurately modeling the intricate and subtle dynamics of…

Computer Vision and Pattern Recognition · Computer Science 2023-11-14 Guinan Su , Yanwu Yang , Zhifeng Li