English
Related papers

Related papers: FAIVConf: Face enhancement for AI-based Video Conf…

200 papers

Almost all digital videos are coded into compact representations before being transmitted. Such compact representations need to be decoded back to pixels before being displayed to humans and - as usual - before being enhanced/analyzed by…

Image and Video Processing · Electrical Eng. & Systems 2023-11-03 Xihua Sheng , Li Li , Dong Liu , Houqiang Li

Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Hanyue Lou , Jinxiu Liang , Minggui Teng , Yi Wang , Boxin Shi

Existing video large language models (VLLMs) primarily leverage prompt agnostic visual encoders, which extract untargeted facial representations without awareness of the queried information, leading to the loss of task critical cues. To…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 Fufangchen Zhao , Songbai Tan , Xuerui Qiu , Linrui Xun , Wenhao Jiang , Jinkai Zheng , Hehe Fan , Jian Gao , Danfeng Yan , Ming Li

Talking-head video editing aims to efficiently insert, delete, and substitute the word of a pre-recorded video through a text transcript editor. The key challenge for this task is obtaining an editing model that generates new talking-head…

Multimedia · Computer Science 2023-09-21 Songlin Yang , Wei Wang , Jun Ling , Bo Peng , Xu Tan , Jing Dong

Although significant progress has been made in face recognition, demographic bias still exists in face recognition systems. For instance, it usually happens that the face recognition performance for a certain demographic group is lower than…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Fu-En Wang , Chien-Yi Wang , Min Sun , Shang-Hong Lai

Talking head video generation aims to animate a human face in a still image with dynamic poses and expressions using motion information derived from a target-driving video, while maintaining the person's identity in the source image.…

Computer Vision and Pattern Recognition · Computer Science 2023-08-21 Fa-Ting Hong , Dan Xu

Facial expression recognition is a challenging classification task that holds broad application prospects in the field of human-computer interaction. This paper aims to introduce the method we will adopt in the 8th Affective and Behavioral…

Computer Vision and Pattern Recognition · Computer Science 2025-03-24 Jun Yu , Yang Zheng , Lei Wang , Yongqi Wang , Shengfan Xu

Video represents the majority of internet traffic today, driving a continual race between the generation of higher quality content, transmission of larger file sizes, and the development of network infrastructure. In addition, the recent…

Image and Video Processing · Electrical Eng. & Systems 2022-04-05 Pulkit Tandon , Shubham Chandak , Pat Pataranutaporn , Yimeng Liu , Anesu M. Mapuranga , Pattie Maes , Tsachy Weissman , Misha Sra

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

Computer Vision and Pattern Recognition · Computer Science 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

Every generation of mobile devices strives to capture video at higher resolution and frame rate than previous ones. This quality increase also requires additional power and computation to capture and encode high-quality media. We propose a…

Image and Video Processing · Electrical Eng. & Systems 2025-03-31 Hidekazu Takahashi , Takefumi Nagumo , Kensei Jo , Aumiller Andreas , Saeed Rad , Rodrigo Caye Daudt , Yoshitaka Miyatani , Hayato Wakabayashi , Christian Brandli

Video Multimethod Assessment Fusion (VMAF) [1], [2], [3] is a popular tool in the industry for measuring coded video quality. In this study, we propose an auditory-inspired frontend in existing VMAF for creating videos of reference and…

Audio and Speech Processing · Electrical Eng. & Systems 2023-08-08 Arijit Biswas , Harald Mundt

Combining face swapping with lip synchronization technology offers a cost-effective solution for customized talking face generation. However, directly cascading existing models together tends to introduce significant interference between…

Computer Vision and Pattern Recognition · Computer Science 2024-05-10 Zeren Zhang , Haibo Qin , Jiayu Huang , Yixin Li , Hui Lin , Yitao Duan , Jinwen Ma

The problem of faces detection in images or video streams is a classical problem of computer vision. The multiple solutions of this problem have been proposed, but the question of their optimality is still open. Many algorithms achieve a…

Computer Vision and Pattern Recognition · Computer Science 2015-11-24 Ilya Kalinovskii , Vladimir Spitsyn

In this paper, we explore an interesting question of what can be obtained from an $8\times8$ pixel video sequence. Surprisingly, it turns out to be quite a lot. We show that when we process this $8\times8$ video with the right set of audio…

Computer Vision and Pattern Recognition · Computer Science 2022-08-18 Sindhu B Hegde , Rudrabha Mukhopadhyay , Vinay P Namboodiri , C. V. Jawahar

To unlock video chat for hundreds of millions of people hindered by poor connectivity or unaffordable data costs, we propose to authentically reconstruct faces on the receiver's device using facial landmarks extracted at the sender's side…

Computer Vision and Pattern Recognition · Computer Science 2020-12-02 Maxime Oquab , Pierre Stock , Oran Gafni , Daniel Haziza , Tao Xu , Peizhao Zhang , Onur Celebi , Yana Hasson , Patrick Labatut , Bobo Bose-Kolanu , Thibault Peyronel , Camille Couprie

We propose an efficient framework, called Simple Swap (SimSwap), aiming for generalized and high fidelity face swapping. In contrast to previous approaches that either lack the ability to generalize to arbitrary identity or fail to preserve…

Computer Vision and Pattern Recognition · Computer Science 2021-06-14 Renwang Chen , Xuanhong Chen , Bingbing Ni , Yanhao Ge

Video face swapping is crucial in film and entertainment production, where achieving high fidelity and temporal consistency over long and complex video sequences remains a significant challenge. Inspired by recent advances in…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Zekai Luo , Zongze Du , Zhouhang Zhu , Hao Zhong , Muzhi Zhu , Wen Wang , Yuling Xi , Chenchen Jing , Hao Chen , Chunhua Shen

Facial video editing has become increasingly important for content creators, enabling the manipulation of facial expressions and attributes. However, existing models encounter challenges such as poor editing quality, high computational…

Computer Vision and Pattern Recognition · Computer Science 2025-01-14 Tharun Anand , Aryan Garg , Kaushik Mitra

For $360^{\circ}$ video streaming, FoV-adaptive coding that allocates more bits for the predicted user's field of view (FoV) is an effective way to maximize the rendered video quality under the limited bandwidth. We develop a low-latency…

Image and Video Processing · Electrical Eng. & Systems 2024-03-19 Yixiang Mao , Liyang Sun , Yong Liu , Yao Wang

In Image-to-Video (I2V) generation, a video is created using an input image as the first-frame condition. Existing I2V methods concatenate the full information of the conditional image with noisy latents to achieve high fidelity. However,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Yunyang Ge , Xinhua Cheng , Chengshu Zhao , Xianyi He , Shenghai Yuan , Bin Lin , Bin Zhu , Li Yuan