English
Related papers

Related papers: Audio-Assisted Face Video Restoration with Tempora…

200 papers

Neural representations for video (NeRV) have gained considerable attention for their strong performance across various video tasks. However, existing NeRV methods often struggle to capture fine spatial details, resulting in vague…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Li Yu , Zhihui Li , Chao Yao , Jimin Xiao , Moncef Gabbouj

The large domain discrepancy between faces captured in polarimetric (or conventional) thermal and visible domain makes cross-domain face recognition quite a challenging problem for both human-examiners and computer vision algorithms.…

Computer Vision and Pattern Recognition · Computer Science 2017-08-10 He Zhang , Vishal M. Patel , Benjamin S. Riggan , Shuowen Hu

Graphics Interchange Format (GIF) is a highly portable graphics format that is ubiquitous on the Internet. Despite their small sizes, GIF images often contain undesirable visual artifacts such as flat color regions, false contours, color…

Computer Vision and Pattern Recognition · Computer Science 2019-04-10 Yang Wang , Haibin Huang , Chuan Wang , Tong He , Jue Wang , Minh Hoai

Video autoencoders compress videos into compact latent representations for efficient reconstruction, playing a vital role in enhancing the quality and efficiency of video generation. However, existing video autoencoders often entangle…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Cuifeng Shen , Lumin Xu , Xingguo Zhu , Gengdai Liu

Despite previous success in generating audio-driven talking heads, most of the previous studies focus on the correlation between speech content and the mouth shape. Facial emotion, which is one of the most important features on natural…

Computer Vision and Pattern Recognition · Computer Science 2021-05-21 Xinya Ji , Hang Zhou , Kaisiyuan Wang , Wayne Wu , Chen Change Loy , Xun Cao , Feng Xu

We present to recover the complete 3D facial geometry from a single depth view by proposing an Attention Guided Generative Adversarial Networks (AGGAN). In contrast to existing work which normally requires two or more depth views to recover…

Computer Vision and Pattern Recognition · Computer Science 2020-09-03 Xiaoxu Cai , Hui Yu , Jianwen Lou , Xuguang Zhang , Gongfa Li , Junyu Dong

Pose-invariant face recognition refers to the problem of identifying or verifying a person by analyzing face images captured from different poses. This problem is challenging due to the large variation of pose, illumination and facial…

Computer Vision and Pattern Recognition · Computer Science 2020-11-11 In Seop Na , Chung Tran , Dung Nguyen , Sang Dinh

Audio is essential for multimodal video understanding. On the one hand, video inherently contains audio, which supplies complementary information to vision. Besides, video large language models (Video-LLMs) can encounter many audio-centric…

Computer Vision and Pattern Recognition · Computer Science 2025-04-04 Yuxin Guo , Shuailei Ma , Shijie Ma , Xiaoyi Bao , Chen-Wei Xie , Kecheng Zheng , Tingyu Weng , Siyang Sun , Yun Zheng , Wei Zou

Generating photo-realistic video portrait with arbitrary speech audio is a crucial problem in film-making and virtual reality. Recently, several works explore the usage of neural radiance field in this task to improve 3D realness and image…

Computer Vision and Pattern Recognition · Computer Science 2023-02-01 Zhenhui Ye , Ziyue Jiang , Yi Ren , Jinglin Liu , JinZheng He , Zhou Zhao

Masked-based autoregressive models have demonstrated promising image generation capability in continuous space. However, their potential for video generation remains under-explored. In this paper, we propose \textbf{VideoMAR}, a concise and…

Computer Vision and Pattern Recognition · Computer Science 2025-06-19 Hu Yu , Biao Gong , Hangjie Yuan , DanDan Zheng , Weilong Chai , Jingdong Chen , Kecheng Zheng , Feng Zhao

Multi-view frame reconstruction is an important problem particularly when multiple frames are missing and past and future frames within the camera are far apart from the missing ones. Realistic coherent frames can still be reconstructed…

Computer Vision and Pattern Recognition · Computer Science 2019-03-05 Tahmida Mahmud , Mohammad Billah , Amit K. Roy-Chowdhury

Humans can easily imagine a scene from auditory information based on their prior knowledge of audio-visual events. In this paper, we mimic this innate human ability in deep learning models to improve the quality of video inpainting. To…

Audio and Speech Processing · Electrical Eng. & Systems 2023-10-12 Kyuyeon Kim , Junsik Jung , Woo Jae Kim , Sung-Eui Yoon

Animating human face images aims to synthesize a desired source identity in a natural-looking way mimicking a driving video's facial movements. In this context, Generative Adversarial Networks have demonstrated remarkable potential in…

Computer Vision and Pattern Recognition · Computer Science 2024-08-26 Alireza Javanmardi , Alain Pagani , Didier Stricker

This paper presents a new problem of unpaired face translation between images and videos, which can be applied to facial video prediction and enhancement. In this problem there exist two major technical challenges: 1) designing a robust…

Computer Vision and Pattern Recognition · Computer Science 2017-12-05 Zhiwu Huang , Bernhard Kratzwald , Danda Pani Paudel , Jiqing Wu , Luc Van Gool

In this paper, we propose a novel attribute-guided cross-resolution (low-resolution to high-resolution) face recognition framework that leverages a coupled generative adversarial network (GAN) structure with adversarial training to find the…

Computer Vision and Pattern Recognition · Computer Science 2019-08-07 Veeru Talreja , Fariborz Taherkhani , Matthew C Valenti , Nasser M Nasrabadi

Neural Representations for Videos (NeRV) has emerged as a promising implicit neural representation (INR) approach for video analysis, which represents videos as neural networks with frame indexes as inputs. However, NeRV-based methods are…

Computer Vision and Pattern Recognition · Computer Science 2025-01-20 Jialong Guo , Ke liu , Jiangchao Yao , Zhihua Wang , Jiajun Bu , Haishuai Wang

Modern text-to-video (T2V) diffusion models can synthesize visually compelling clips, yet they remain brittle at fine-scale structure: even state-of-the-art generators often produce distorted faces and hands, warped backgrounds, and…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Tejas Panambur , Ishan Rajendrakumar Dave , Chongjian Ge , Ersin Yumer , Xue Bai

With the rapid advancement of sophisticated synthetic audio-visual content, e.g., for subtle malicious manipulations, ensuring the integrity of digital media has become paramount. This work presents a novel approach to temporal localization…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Christos Koutlis , Symeon Papadopoulos

Beyond the existing single-person and multiple-person human parsing tasks in static images, this paper makes the first attempt to investigate a more realistic video instance-level human parsing that simultaneously segments out each person…

Computer Vision and Pattern Recognition · Computer Science 2018-08-13 Qixian Zhou , Xiaodan Liang , Ke Gong , Liang Lin

Real-world talking faces often accompany with natural head movement. However, most existing talking face video generation methods only consider facial animation with fixed head pose. In this paper, we address this problem by proposing a…

Computer Vision and Pattern Recognition · Computer Science 2020-03-06 Ran Yi , Zipeng Ye , Juyong Zhang , Hujun Bao , Yong-Jin Liu