English
Related papers

Related papers: Assessing Identity Leakage in Talking Face Generat…

200 papers

Speech-driven animation has gained significant traction in recent years, with current methods achieving near-photorealistic results. However, the field remains underexplored regarding non-verbal communication despite evidence demonstrating…

Computer Vision and Pattern Recognition · Computer Science 2023-08-31 Antoni Bigata Casademunt , Rodrigo Mira , Nikita Drobyshev , Konstantinos Vougioukas , Stavros Petridis , Maja Pantic

Recent works on language-guided image manipulation have shown great power of language in providing rich semantics, especially for face images. However, the other natural information, motions, in language is less explored. In this paper, we…

Computer Vision and Pattern Recognition · Computer Science 2024-07-04 Tiankai Hang , Huan Yang , Bei Liu , Jianlong Fu , Xin Geng , Baining Guo

We propose a two-stage framework for audio-driven talking head generation with fine-grained expression control via facial Action Units (AUs). Unlike prior methods relying on emotion labels or implicit AU conditioning, our model explicitly…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Shao-Yu Chang , Jingyi Xu , Hieu Le , Dimitris Samaras

Biometrics authentication has become increasingly popular due to its security and convenience; however, traditional biometrics are becoming less desirable in scenarios such as new mobile devices, Virtual Reality, and Smart Vehicles. For…

Computer Vision and Pattern Recognition · Computer Science 2025-11-10 Huashan Chen , Yifan Xu , Yue Feng , Ming Jian , Feng Liu , Pengfei Hu , Kebin Peng , Sen He , Zi Wang

Micro-expressions have drawn increasing interest lately due to various potential applications. The task is, however, difficult as it incorporates many challenges from the fields of computer vision, machine learning and emotional sciences.…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Tuomas Varanka , Yante Li , Wei Peng , Guoying Zhao

Generative adversarial networks (GANs) are able to generate high resolution photo-realistic images of objects that "do not exist." These synthetic images are rather difficult to detect as fake. However, the manner in which these generative…

Computer Vision and Pattern Recognition · Computer Science 2021-01-14 Patrick Tinsley , Adam Czajka , Patrick Flynn

Generating realistic lip motion from audio to simulate speech production is critical for driving natural character animation. Previous research has shown that traditional metrics used to optimize and assess models for generating lip motion…

Lip reading aims to predict speech based on lip movements alone. As it focuses on visual information to model the speech, its performance is inherently sensitive to personal lip appearances and movements. This makes the lip reading models…

Computer Vision and Pattern Recognition · Computer Science 2022-08-10 Minsu Kim , Hyunjun Kim , Yong Man Ro

Generative spoken language models produce speech in a wide range of voices, prosody, and recording conditions, seemingly approaching the diversity of natural speech. However, the extent to which generated speech is acoustically diverse…

Audio and Speech Processing · Electrical Eng. & Systems 2025-03-12 Matthieu Futeral , Andrea Agostinelli , Marco Tagliasacchi , Neil Zeghidour , Eugene Kharitonov

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

We introduce GenSync, a novel framework for multi-identity lip-synced video synthesis using 3D Gaussian Splatting. Unlike most existing 3D methods that require training a new model for each identity , GenSync learns a unified network that…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Anushka Agarwal , Muhammad Yusuf Hassan , Talha Chafekar

We propose a novel method for generating high-resolution videos of talking-heads from speech audio and a single 'identity' image. Our method is based on a convolutional neural network model that incorporates a pre-trained StyleGAN…

Computer Vision and Pattern Recognition · Computer Science 2022-09-12 Mohammed M. Alghamdi , He Wang , Andrew J. Bulpitt , David C. Hogg

Large Language Model (LLM) outputs often vary across user sociodemographic attributes, leading to disparities in factual accuracy, utility, and safety, even for objective questions where demographic information is irrelevant. Unlike prior…

Computation and Language · Computer Science 2026-01-15 Miao Zhang , Kelly Chen , Md Mehrab Tanjim , Rumi Chunara

In this work, we propose an ID-preserving talking head generation framework, which advances previous methods in two aspects. First, as opposed to interpolating from sparse flow, we claim that dense landmarks are crucial to achieving…

Computer Vision and Pattern Recognition · Computer Science 2023-03-28 Bowen Zhang , Chenyang Qi , Pan Zhang , Bo Zhang , HsiangTao Wu , Dong Chen , Qifeng Chen , Yong Wang , Fang Wen

Human emotional expression is inherently dynamic, complex, and fluid, characterized by smooth transitions in intensity throughout verbal communication. However, the modeling of such intensity fluctuations has been largely overlooked by…

Sound · Computer Science 2024-10-01 Jingyi Xu , Hieu Le , Zhixin Shu , Yang Wang , Yi-Hsuan Tsai , Dimitris Samaras

Creating a realistic animatable avatar from a single static portrait remains challenging. Existing approaches often struggle to capture subtle facial expressions, the associated global body movements, and the dynamic background. To address…

Computer Vision and Pattern Recognition · Computer Science 2025-04-08 Mengchao Wang , Qiang Wang , Fan Jiang , Yaqi Fan , Yunpeng Zhang , Yonggang Qi , Kun Zhao , Mu Xu

A major challenge in DeepFake forgery detection is that state-of-the-art algorithms are mostly trained to detect a specific fake method. As a result, these approaches show poor generalization across different types of facial manipulations,…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Davide Cozzolino , Andreas Rössler , Justus Thies , Matthias Nießner , Luisa Verdoliva

People talk with diversified styles. For one piece of speech, different talking styles exhibit significant differences in the facial and head pose movements. For example, the "excited" style usually talks with the mouth wide open, while the…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Haozhe Wu , Jia Jia , Haoyu Wang , Yishun Dou , Chao Duan , Qingshan Deng

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Jinting Wang , Jun Wang , Hei Victor Cheng , Li Liu

With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets, resulting in high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Hao Wu , Xiangyang Luo , Hao Wang , Jiawei Zhang , Yi Zhang , Jinwei Wang