English
Related papers

Related papers: A Unified Compression Framework for Efficient Spee…

200 papers

The challenge of talking face generation from speech lies in aligning two different modal information, audio and video, such that the mouth region corresponds to input audio. Previous methods either exploit audio-visual representation…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Se Jin Park , Minsu Kim , Joanna Hong , Jeongsoo Choi , Yong Man Ro

Unlike existing methods that rely on source images as appearance references and use source speech to generate motion, this work proposes a novel approach that directly extracts information from the speech, addressing key challenges in…

Audio and Speech Processing · Electrical Eng. & Systems 2026-03-03 Jinting Wang , Jun Wang , Hei Victor Cheng , Li Liu

3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video that lacks the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Junuk Cha , Seongro Yoon , Valeriya Strizhkova , Francois Bremond , Seungryul Baek

Talking face generation has historically struggled to produce head movements and natural facial expressions without guidance from additional reference videos. Recent developments in diffusion-based generative models allow for more realistic…

Computer Vision and Pattern Recognition · Computer Science 2023-08-01 Michał Stypułkowski , Konstantinos Vougioukas , Sen He , Maciej Zięba , Stavros Petridis , Maja Pantic

Talking face generation, also known as speech-to-lip generation, reconstructs facial motions concerning lips given coherent speech input. The previous studies revealed the importance of lip-speech synchronization and visual quality. Despite…

Computer Vision and Pattern Recognition · Computer Science 2023-03-31 Jiadong Wang , Xinyuan Qian , Malu Zhang , Robby T. Tan , Haizhou Li

To improve the experiences of face-to-face conversation with avatar, this paper presents a novel conversation system. It is composed of two sequence-to-sequence models respectively for listening and speaking and a Generative Adversarial…

Computer Vision and Pattern Recognition · Computer Science 2019-08-22 Zezhou Chen , Zhaoxiang Liu , Huan Hu , Jinqiang Bai , Shiguo Lian , Fuyuan Shi , Kai Wang

We present a novel approach for generating realistic speaking and talking faces by synthesizing a person's voice and facial movements from a static image, a voice profile, and a target text. The model encodes the prompt/driving text, the…

Computer Vision and Pattern Recognition · Computer Science 2026-02-24 Aashish Chandra , Aashutosh A , Abhijit Das

With the rapid advancement of diffusion models, talking face generation has made remarkable progress. However, existing diffusion-based methods still require task-specific fine-tuning and large-scale audiovisual datasets, resulting in high…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Hao Wu , Xiangyang Luo , Hao Wang , Jiawei Zhang , Yi Zhang , Jinwei Wang

Audio-driven cospeech video generation typically involves two stages: speech-to-gesture and gesture-to-video. While significant advances have been made in speech-to-gesture generation, synthesizing natural expressions and gestures remains…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Renda Li , Xiaohua Qi , Qiang Ling , Jun Yu , Ziyi Chen , Peng Chang , Mei HanJing Xiao

When people deliver a speech, they naturally move heads, and this rhythmic head motion conveys prosodic information. However, generating a lip-synced video while moving head naturally is challenging. While remarkably successful, existing…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Lele Chen , Guofeng Cui , Celong Liu , Zhong Li , Ziyi Kou , Yi Xu , Chenliang Xu

Generating semantically coherent and visually accurate talking faces requires bridging the gap between linguistic meaning and facial articulation. Although audio-driven methods remain prevalent, their reliance on high-quality paired audio…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Xu Wang , Shengeng Tang , Fei Wang , Lechao Cheng , Dan Guo , Feng Xue , Richang Hong

We present Livatar, a real-time audio-driven talking heads videos generation framework. Existing baselines suffer from limited lip-sync accuracy and long-term pose drift. We address these limitations with a flow matching based framework.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-28 Haiyang Liu , Xiaolin Hong , Xuancheng Yang , Yudi Ruan , Xiang Lian , Michael Lingelbach , Hongwei Yi , Wei Li

The goal of a speech-to-image transform is to produce a photo-realistic picture directly from a speech signal. Recently, various studies have focused on this task and have achieved promising performance. However, current speech-to-image…

Computer Vision and Pattern Recognition · Computer Science 2023-05-18 Zhenxing Zhang , Lambert Schomaker

The state-of-the-art in text-to-speech synthesis has recently improved considerably due to novel neural waveform generation methods, such as WaveNet. However, these methods suffer from their slow sequential inference process, while their…

Audio and Speech Processing · Electrical Eng. & Systems 2018-10-31 Lauri Juvela , Bajibabu Bollepalli , Junichi Yamagishi , Paavo Alku

We propose an audio-driven talking-head method to generate photo-realistic talking-head videos from a single reference image. In this work, we tackle two key challenges: (i) producing natural head motions that match speech prosody, and (ii)…

Computer Vision and Pattern Recognition · Computer Science 2021-07-21 Suzhen Wang , Lincheng Li , Yu Ding , Changjie Fan , Xin Yu

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Bin Liu , Zhixiang Xiong , Zhifen He , Bo Li

This paper presents a simple method for speech videos generation based on audio: given a piece of audio, we can generate a video of the target face speaking this audio. We propose Generative Adversarial Networks (GAN) with cut speech audio…

Sound · Computer Science 2022-07-20 Hanhaodi Zhang

While dynamic Neural Radiance Fields (NeRF) have shown success in high-fidelity 3D modeling of talking portraits, the slow training and inference speed severely obstruct their potential usage. In this paper, we propose an efficient…

Computer Vision and Pattern Recognition · Computer Science 2022-11-23 Jiaxiang Tang , Kaisiyuan Wang , Hang Zhou , Xiaokang Chen , Dongliang He , Tianshu Hu , Jingtuo Liu , Gang Zeng , Jingdong Wang

We propose a neural talking-head video synthesis model and demonstrate its application to video conferencing. Our model learns to synthesize a talking-head video using a source image containing the target person's appearance and a driving…

Computer Vision and Pattern Recognition · Computer Science 2021-04-06 Ting-Chun Wang , Arun Mallya , Ming-Yu Liu

Individuals have unique facial expression and head pose styles that reflect their personalized speaking styles. Existing one-shot talking head methods cannot capture such personalized characteristics and therefore fail to produce diverse…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Suzhen Wang , Yifeng Ma , Yu Ding , Zhipeng Hu , Changjie Fan , Tangjie Lv , Zhidong Deng , Xin Yu
‹ Prev 1 3 4 5 6 7 10 Next ›