English
Related papers

Related papers: Audio-Assisted Face Video Restoration with Tempora…

200 papers

Despite rapid advances in face recognition, there remains a clear gap between the performance of still image-based face recognition and video-based face recognition, due to the vast difference in visual quality between the domains and the…

Computer Vision and Pattern Recognition · Computer Science 2017-08-15 Kihyuk Sohn , Sifei Liu , Guangyu Zhong , Xiang Yu , Ming-Hsuan Yang , Manmohan Chandraker

We propose in this paper a new paradigm for facial video compression. We leverage the generative capacity of GANs such as StyleGAN to represent and compress a video, including intra and inter compression. Each frame is inverted in the…

Image and Video Processing · Electrical Eng. & Systems 2022-07-14 Mustafa Shukor , Bharath Bhushan Damodaran , Xu Yao , Pierre Hellier

High-fidelity facial avatar reconstruction from a monocular video is a significant research problem in computer graphics and computer vision. Recently, Neural Radiance Field (NeRF) has shown impressive novel view rendering results and has…

Computer Vision and Pattern Recognition · Computer Science 2023-03-07 Yunpeng Bai , Yanbo Fan , Xuan Wang , Yong Zhang , Jingxiang Sun , Chun Yuan , Ying Shan

Facial expression generation has always been an intriguing task for scientists and researchers all over the globe. In this context, we present our novel approach for generating videos of the six basic facial expressions. Starting from a…

Computer Vision and Pattern Recognition · Computer Science 2023-07-04 Hamza Bouzid , Lahoucine Ballihi

We present RAVEn, a self-supervised multi-modal approach to jointly learn visual and auditory speech representations. Our pre-training objective involves encoding masked inputs, and then predicting contextualised targets generated by…

Machine Learning · Computer Science 2023-04-06 Alexandros Haliassos , Pingchuan Ma , Rodrigo Mira , Stavros Petridis , Maja Pantic

Many recent works have been proposed for face image editing by leveraging the latent space of pretrained GANs. However, few attempts have been made to directly apply them to videos, because 1) they do not guarantee temporal consistency, 2)…

Computer Vision and Pattern Recognition · Computer Science 2022-06-28 Jiyang Yu , Jingen Liu , Jing Huang , Wei Zhang , Tao Mei

Audio-Visual Navigation (AVN) requires an embodied agent to navigate toward a sound source by utilizing both vision and binaural audio. A core challenge arises in complex acoustic environments, where binaural cues become intermittently…

Sound · Computer Science 2026-04-06 Teng Liu , Yinfeng Yu

Recovering a photorealistic face from an artistic portrait is a challenging task since crucial facial details are often distorted or completely lost in artistic compositions. To handle this loss, we propose an Attribute-guided Face Recovery…

Computer Vision and Pattern Recognition · Computer Science 2019-04-09 Fatemeh Shiri , Xin Yu , Fatih Porikli , Richard Hartley , Piotr Koniusz

Audio-Driven Talking Face Generation aims at generating realistic videos of talking faces, focusing on accurate audio-lip synchronization without deteriorating any identity-related visual details. Recent state-of-the-art methods are based…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Dogucan Yaman , Fevziye Irem Eyiokur , Leonard Bärmann , Hazım Kemal Ekenel , Alexander Waibel

We propose a novel self-supervised approach for learning audio and visual representations from unlabeled videos, based on their correspondence. The approach uses an attention mechanism to learn the relative importance of convolutional…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Sudha Krishnamurthy

Person Re-Identification (person re-id) is a crucial task as its applications in visual surveillance and human-computer interaction. In this work, we present a novel joint Spatial and Temporal Attention Pooling Network (ASTPN) for…

Computer Vision and Pattern Recognition · Computer Science 2017-10-02 Shuangjie Xu , Yu Cheng , Kang Gu , Yang Yang , Shiyu Chang , Pan Zhou

Reconstructing static 3D scene from monocular video with dynamic objects is important for numerous applications such as virtual reality and autonomous driving. Current approaches typically rely on background for static scene reconstruction,…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Yedong Shen , Shiqi Zhang , Sha Zhang , Yifan Duan , Xinran Zhang , Wenhao Yu , Lu Zhang , Jiajun Deng , Yanyong Zhang

Generating talking person portraits with arbitrary speech audio is a crucial problem in the field of digital human and metaverse. A modern talking face generation method is expected to achieve the goals of generalized audio-lip…

Computer Vision and Pattern Recognition · Computer Science 2023-05-02 Zhenhui Ye , Jinzheng He , Ziyue Jiang , Rongjie Huang , Jiawei Huang , Jinglin Liu , Yi Ren , Xiang Yin , Zejun Ma , Zhou Zhao

The current research focus on Content-Based Video Retrieval requires higher-level video representation describing the long-range semantic dependencies of relevant incidents, events, etc. However, existing methods commonly process the frames…

Computer Vision and Pattern Recognition · Computer Science 2020-10-01 Jie Shao , Xin Wen , Bingchen Zhao , Xiangyang Xue

Deep-Learning-based video recognition has shown promising improvements along with the development of large-scale datasets and spatiotemporal network architectures. In image recognition, learning spatially invariant features is a key factor…

Computer Vision and Pattern Recognition · Computer Science 2020-08-14 Taeoh Kim , Hyeongmin Lee , MyeongAh Cho , Ho Seong Lee , Dong Heon Cho , Sangyoun Lee

tmospheric turbulence presents a significant challenge in long-range imaging. Current restoration algorithms often struggle with temporal inconsistency, as well as limited generalization ability across varying turbulence levels and scene…

Image and Video Processing · Electrical Eng. & Systems 2023-12-11 Haoming Cai , Jingxi Chen , Brandon Y. Feng , Weiyun Jiang , Mingyang Xie , Kevin Zhang , Ashok Veeraraghavan , Christopher Metzler

Endoscopic videos from multicentres often have different imaging conditions, e.g., color and illumination, which make the models trained on one domain usually fail to generalize well to another. Domain adaptation is one of the potential…

Computer Vision and Pattern Recognition · Computer Science 2020-04-20 Jiawei Chen , Yuexiang Li , Kai Ma , Yefeng Zheng

Recent advancements in self-supervised audio-visual representation learning have demonstrated its potential to capture rich and comprehensive representations. However, despite the advantages of data augmentation verified in many learning…

Machine Learning · Computer Science 2024-06-21 Jongsuk Kim , Hyeongkeun Lee , Kyeongha Rho , Junmo Kim , Joon Son Chung

DeepFake based digital facial forgery is threatening the public media security, especially when lip manipulation has been used in talking face generation, the difficulty of fake video detection is further improved. By only changing lip…

Computer Vision and Pattern Recognition · Computer Science 2022-03-11 Ganglai Wang , Peng Zhang , Lei Xie , Wei Huang , Yufei Zha , Yanning Zhang

In this paper, we propose to detect forged videos, of faces, in online videos. To facilitate this detection, we propose to use smaller (fewer parameters to learn) convolutional neural networks (CNN), for a data-driven approach to forged…

Computer Vision and Pattern Recognition · Computer Science 2020-05-19 Neilesh Sambhu , Shaun Canavan