English
Related papers

Related papers: FlashLips: 100-FPS Mask-Free Latent Lip-Sync using…

200 papers

Editing real facial images is a crucial task in computer vision with significant demand in various real-world applications. While GAN-based methods have showed potential in manipulating images especially when combined with CLIP, these…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Dongxu Yue , Qin Guo , Munan Ning , Jiaxi Cui , Yuesheng Zhu , Li Yuan

We present an efficient text-to-video generation framework based on latent diffusion models, termed MagicVideo. MagicVideo can generate smooth video clips that are concordant with the given text descriptions. Due to a novel and efficient 3D…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Daquan Zhou , Weimin Wang , Hanshu Yan , Weiwei Lv , Yizhe Zhu , Jiashi Feng

Generating realistic talking-head videos remains challenging due to persistent issues such as imperfect lip synchronization, unnatural motion, and evaluation metrics that correlate poorly with human perception. We propose FlowPortrait, a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Weiting Tan , Andy T. Liu , Ming Tu , Xinghua Qu , Philipp Koehn , Lu Lu

Recent works on audio-driven talking head synthesis using Neural Radiance Fields (NeRF) have achieved impressive results. However, due to inadequate pose and expression control caused by NeRF implicit representation, these methods still…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Hongyun Yu , Zhan Qu , Qihang Yu , Jianchuan Chen , Zhonghua Jiang , Zhiwen Chen , Shengyu Zhang , Jimin Xu , Fei Wu , Chengfei Lv , Gang Yu

Recent advancement in Generative Adversarial Networks in speech synthesis domain[3],[2] have shown, that it's possible to train GANs [8] in a reliable manner for high quality coherent waveform generation from mel-spectograms. We propose…

Audio and Speech Processing · Electrical Eng. & Systems 2020-06-16 Luka Chkhetiani , Levan Bejanidze

The presence of a corresponding talking face has been shown to significantly improve speech intelligibility in noisy conditions and for hearing impaired population. In this paper, we present a system that can generate landmark points of a…

Computer Vision and Pattern Recognition · Computer Science 2018-04-24 Sefik Emre Eskimez , Ross K Maddox , Chenliang Xu , Zhiyao Duan

Sign language production (SLP) aims to translate spoken language sentences into a sequence of pose frames in a sign language, bridging the communication gap and promoting digital inclusion for deaf and hard-of-hearing communities. Existing…

Computation and Language · Computer Science 2025-09-16 Liqian Feng , Lintao Wang , Kun Hu , Dehui Kong , Zhiyong Wang

The fast evolution and widespread of deepfake techniques in real-world scenarios require stronger generalization abilities of face forgery detectors. Some works capture the features that are unrelated to method-specific artifacts, such as…

Computer Vision and Pattern Recognition · Computer Science 2022-03-03 Hanqing Zhao , Wenbo Zhou , Dongdong Chen , Weiming Zhang , Nenghai Yu

Diffusion models are a class of generative models that have been recently used for speech enhancement with remarkable success but are computationally expensive at inference time. Therefore, these models are impractical for processing…

Audio and Speech Processing · Electrical Eng. & Systems 2025-09-15 Bunlong Lay , Rostislav Makarov , Timo Gerkmann

Transformer is leading a trend in the field of image processing. Despite the great success that existing lightweight image processing transformers have achieved, they are tailored to FLOPs or parameters reduction, rather than practical…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Junbo Qiao , Wei Li , Haizhen Xie , Hanting Chen , Yunshuai Zhou , Zhijun Tu , Jie Hu , Shaohui Lin

The recent success of text-to-image synthesis has taken the world by storm and captured the general public's imagination. From a technical standpoint, it also marked a drastic change in the favored architecture to design generative image…

Computer Vision and Pattern Recognition · Computer Science 2023-06-21 Minguk Kang , Jun-Yan Zhu , Richard Zhang , Jaesik Park , Eli Shechtman , Sylvain Paris , Taesung Park

Diffusion language models offer parallel token generation and inherent bidirectionality, promising more efficient and powerful sequence modeling compared to autoregressive approaches. However, state-of-the-art diffusion models (e.g., Dream…

Computation and Language · Computer Science 2025-10-10 Zhanqiu Hu , Jian Meng , Yash Akhauri , Mohamed S. Abdelfattah , Jae-sun Seo , Zhiru Zhang , Udit Gupta

Large language model (LLM)-based text-to-speech (TTS) systems achieve remarkable naturalness via autoregressive (AR) decoding, but require N sequential steps to generate N speech tokens. We present LLaDA-TTS, which replaces the AR LLM with…

Sound · Computer Science 2026-03-30 Xiaoyu Fan , Huizhi Xie , Wei Zou , Yunzhang Chen

We show that pre-trained Generative Adversarial Networks (GANs) such as StyleGAN and BigGAN can be used as a latent bank to improve the performance of image super-resolution. While most existing perceptual-oriented approaches attempt to…

Computer Vision and Pattern Recognition · Computer Science 2022-08-01 Kelvin C. K. Chan , Xiangyu Xu , Xintao Wang , Jinwei Gu , Chen Change Loy

We propose LatentSwap, a simple face swapping framework generating a face swap latent code of a given generator. Utilizing randomly sampled latent codes, our framework is light and does not require datasets besides employing the pre-trained…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Changho Choi , Minho Kim , Junhyeok Lee , Hyoung-Kyu Song , Younggeun Kim , Seungryong Kim

Deep generative models, like GANs, have considerably improved the state of the art in image synthesis, and are able to generate near photo-realistic images in structured domains such as human faces. Based on this success, recent work on…

Computer Vision and Pattern Recognition · Computer Science 2022-03-10 Guillaume Couairon , Asya Grechka , Jakob Verbeek , Holger Schwenk , Matthieu Cord

3D Gaussian splatting-based talking head synthesis has recently gained attention for its ability to render high-fidelity images with real-time inference speed. However, since it is typically trained on only a short video that lacks the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-04 Junuk Cha , Seongro Yoon , Valeriya Strizhkova , Francois Bremond , Seungryul Baek

Lip sync is a fundamental audio-visual task. However, existing lip sync methods fall short of being robust in the wild. One important cause could be distracting factors on the visual input side, making extracting lip motion information…

Computer Vision and Pattern Recognition · Computer Science 2024-09-10 Chun Wang

Audio-Visual Speech-to-Speech Translation typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony-ensuring that the movements of the lips match the…

Map-free LiDAR localization systems accurately localize within known environments by predicting sensor position and orientation directly from raw point clouds, eliminating the need for large maps and descriptors. However, their long…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Raktim Gautam Goswami , Naman Patel , Prashanth Krishnamurthy , Farshad Khorrami