English
Related papers

Related papers: UniTalking: A Unified Audio-Video Framework for Ta…

200 papers

Talking head generation with arbitrary identities and speech audio remains a crucial problem in the realm of the virtual metaverse. Recently, diffusion models have become a popular generative technique in this field with their strong…

Graphics · Computer Science 2025-08-11 Xinyang Li , Gen Li , Zhihui Lin , Yichen Qian , GongXin Yao , Weinan Jia , Aowen Wang , Weihua Chen , Fan Wang

Multimodal-driven talking face generation refers to animating a portrait with the given pose, expression, and gaze transferred from the driving image and video, or estimated from the text and audio. However, existing methods ignore the…

Computer Vision and Pattern Recognition · Computer Science 2023-05-10 Chao Xu , Shaoting Zhu , Junwei Zhu , Tianxin Huang , Jiangning Zhang , Ying Tai , Yong Liu

Real-time video dubbing that preserves identity consistency while achieving accurate lip synchronization remains a critical challenge. Existing approaches face a trilemma: diffusion-based methods achieve high visual fidelity but suffer from…

Computer Vision and Pattern Recognition · Computer Science 2025-03-27 Yue Zhang , Zhizhou Zhong , Minhao Liu , Zhaokang Chen , Bin Wu , Yubin Zeng , Chao Zhan , Yingjie He , Junxin Huang , Wenjiang Zhou

Speech-driven 3D talking head generation aims to produce lifelike facial animations precisely synchronized with speech. While considerable progress has been made in achieving high lip-synchronization accuracy, existing methods largely…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Bin Wang , Yang Xu , Huan Zhao , Hao Zhang , Zixing Zhang

Generating talking face videos from audio attracts lots of research interest. A few person-specific methods can generate vivid videos but require the target speaker's videos for training or fine-tuning. Existing person-generic methods have…

Computer Vision and Pattern Recognition · Computer Science 2023-05-16 Weizhi Zhong , Chaowei Fang , Yinqi Cai , Pengxu Wei , Gangming Zhao , Liang Lin , Guanbin Li

Recent studies in speech-driven 3D talking head generation have achieved convincing results in verbal articulations. However, generating accurate lip-syncs degrades when applied to input speech in other languages, possibly due to the lack…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Kim Sung-Bin , Lee Chae-Yeon , Gihun Son , Oh Hyun-Bin , Janghoon Ju , Suekyeong Nam , Tae-Hyun Oh

Building an end-to-end conversational agent for multi-domain task-oriented dialogues has been an open challenge for two main reasons. First, tracking dialogue states of multiple domains is non-trivial as the dialogue agent must obtain…

Computation and Language · Computer Science 2020-11-17 Hung Le , Doyen Sahoo , Chenghao Liu , Nancy F. Chen , Steven C. H. Hoi

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

Computer Vision and Pattern Recognition · Computer Science 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu

While audio-visual speech models can yield superior performance and robustness compared to audio-only models, their development and adoption are hindered by the lack of labeled and unlabeled audio-visual data and the cost to deploy one…

Computation and Language · Computer Science 2022-11-29 Wei-Ning Hsu , Bowen Shi

In this paper, we propose Universal Holistic Audio Generation (UniHAGen), a task for synthesizing comprehensive auditory scenes that include both on-screen and off-screen sounds across diverse domains (e.g., ambient events, musical…

Sound · Computer Science 2026-04-07 Weiguo Pian , Saksham Singh Kushwaha , Zhimin Chen , Shijian Deng , Kai Wang , Yunhui Guo , Yapeng Tian

Talking Head Generation aims at synthesizing natural-looking talking videos from speech and a single portrait image. Previous 3D talking head generation methods have relied on domain-specific heuristics such as warping-based facial motion…

Computer Vision and Pattern Recognition · Computer Science 2026-01-27 Tong Shi , Melonie de Almeida , Daniela Ivanova , Nicolas Pugeault , Paul Henderson

Achieving high-fidelity lip-speech synchronization in audio-driven talking portrait synthesis remains challenging. While multi-stage pipelines or diffusion models yield high-quality results, they suffer from high computational costs. Some…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Ziqi Ni , Ao Fu , Yi Zhou

Long-duration talking video synthesis faces enduring challenges in achieving high video quality, portrait consistency, temporal coherence, and computational efficiency. As video length increases, issues such as visual degradation, portrait…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Haojie Zhang , Zhihao Liang , Ruibo Fu , Bingyan Liu , Zhengqi Wen , Xuefei Liu , Jianhua Tao , Yaling Liang

We propose Kling-Foley, a large-scale multimodal Video-to-Audio generation model that synthesizes high-quality audio synchronized with video content. In Kling-Foley, we introduce multimodal diffusion transformers to model the interactions…

We introduce GenSync, a novel framework for multi-identity lip-synced video synthesis using 3D Gaussian Splatting. Unlike most existing 3D methods that require training a new model for each identity , GenSync learns a unified network that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Anushka Agarwal , Muhammad Yusuf Hassan , Talha Chafekar

With the rapid advancement of diffusion-based generative models, portrait image animation has achieved remarkable results. However, it still faces challenges in temporally consistent video generation and fast sampling due to its iterative…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Taekyung Ki , Dongchan Min , Gyeongsu Chae

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu

Audio-driven talking head generation necessitates seamless integration of audio and visual data amidst the challenges posed by diverse input portraits and intricate correlations between audio and facial motions. In response, we propose a…

Computer Vision and Pattern Recognition · Computer Science 2024-12-16 Ziqi Zhou , Weize Quan , Hailin Shi , Wei Li , Lili Wang , Dong-Ming Yan

Although existing unified models achieve strong performance in vision-language understanding and text-to-image generation, they remain limited in addressing image perception and manipulation -- capabilities increasingly demanded in…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Bin Lin , Zongjian Li , Xinhua Cheng , Yuwei Niu , Yang Ye , Xianyi He , Shenghai Yuan , Wangbo Yu , Shaodong Wang , Yunyang Ge , Yatian Pang , Li Yuan

Audio-driven talking head generation holds significant potential for film production. While existing 3D methods have advanced motion modeling and content synthesis, they often produce rendering artifacts, such as motion blur, temporal…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Kui Jiang , Shiyu Liu , Junjun Jiang , Hongxun Yao , Xiaopeng Fan