中文
相关论文

相关论文: Low Bandwidth Video-Chat Compression using Deep Ge…

200 篇论文

In this paper, we explore an interesting question of what can be obtained from an $8\times8$ pixel video sequence. Surprisingly, it turns out to be quite a lot. We show that when we process this $8\times8$ video with the right set of audio…

计算机视觉与模式识别 · 计算机科学 2022-08-18 Sindhu B Hegde , Rudrabha Mukhopadhyay , Vinay P Namboodiri , C. V. Jawahar

We propose a novel end-to-end deep architecture for face landmark detection, based on a deep convolutional and deconvolutional network followed by carefully designed recurrent network structures. The pipeline of this architecture consists…

计算机视觉与模式识别 · 计算机科学 2016-11-01 Hanjiang Lai , Shengtao Xiao , Yan Pan , Zhen Cui , Jiashi Feng , Chunyan Xu , Jian Yin , Shuicheng Yan

Fine-grained classification remains a challenging task because distinguishing categories needs learning complex and local differences. Diversity in the pose, scale, and position of objects in an image makes the problem even more difficult.…

计算机视觉与模式识别 · 计算机科学 2021-09-03 Mahdi Darvish , Mahsa Pouramini , Hamid Bahador

In this work, we propose a novel framework to enable diffusion models to adapt their generation quality based on real-time network bandwidth constraints. Traditional diffusion models produce high-fidelity images by performing a fixed number…

计算机视觉与模式识别 · 计算机科学 2026-04-10 Xi Zhang , Hanwei Zhu , Yan Zhong , Jiamang Wang , Weisi Lin

In this paper, we propose a talking face generation method that takes an audio signal as input and a short target video clip as reference, and synthesizes a photo-realistic video of the target face with natural lip motions, head poses, and…

计算机视觉与模式识别 · 计算机科学 2021-08-19 Chenxu Zhang , Yifan Zhao , Yifei Huang , Ming Zeng , Saifeng Ni , Madhukar Budagavi , Xiaohu Guo

Facial video re-targeting is a challenging problem aiming to modify the facial attributes of a target subject in a seamless manner by a driving monocular sequence. We leverage the 3D geometry of faces and Generative Adversarial Networks…

计算机视觉与模式识别 · 计算机科学 2021-09-29 Michail Christos Doukas , Mohammad Rami Koujan , Viktoriia Sharmanska , Anastasios Roussos

Deep graph generative modeling has proven capable of learning the distribution of complex, multi-scale structures characterizing real-world graphs. However, one of the main limitations of existing methods is their large output space, which…

机器学习 · 计算机科学 2023-06-01 Nathaniel Diamant , Alex M. Tseng , Kangway V. Chuang , Tommaso Biancalani , Gabriele Scalia

Facial landmark detection is a crucial prerequisite for many face analysis applications. Deep learning-based methods currently dominate the approach of addressing the facial landmark detection. However, such works generally introduce a…

计算机视觉与模式识别 · 计算机科学 2019-11-21 Yang Zhao , Yifan Liu , Chunhua Shen , Yongsheng Gao , Shengwu Xiong

Face image super-resolution aims to recover high-resolution facial images from severely degraded inputs. Under extreme upscaling factors, fine facial details are often lost, making accurate reconstruction challenging. Existing methods…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Riccardo Carraro , Anna Briotto , Endi Hysa , Marco Fiorucci , Lamberto Ballan

In the recent years, there has been a significant improvement in the quality of samples produced by (deep) generative models such as variational auto-encoders and generative adversarial networks. However, the representation capabilities of…

图像与视频处理 · 电气工程与系统科学 2026-03-31 Shady Abu Hussein , Tom Tirer , Raja Giryes

Under the limited storage, computing and network bandwidth resources, the video compression coding technology plays an important role for visual communication. To efficiently compress raw video data, a colorization-based video compression…

图像与视频处理 · 电气工程与系统科学 2019-12-24 Zhaoqing Pan , Feng Yuan , Jianjun Lei , Sam Kwong

Mobile and embedded machine learning developers frequently have to compromise between two inferior on-device deployment strategies: sacrifice accuracy and aggressively shrink their models to run on dedicated low-power cores; or sacrifice…

机器学习 · 计算机科学 2023-03-17 Haiguang Li , Trausti Thormundsson , Ivan Poupyrev , Nicholas Gillian

Semantic communication focuses on conveying the task-relevant meaning rather than exact bitwise recovery. For image transmission with a generative receiver, relying only on text descriptions can be insufficient to preserve instance-specific…

信息论 · 计算机科学 2026-01-27 Xuesong Wang , Xinyan Xie , Mo Li , Zhaoqian Liu

Arguably the most common and salient object in daily video communications is the talking head, as encountered in social media, virtual classrooms, teleconferences, news broadcasting, talk shows, etc. When communication bandwidth is limited…

计算机视觉与模式识别 · 计算机科学 2022-03-07 Xi Zhang , Xiaolin Wu

In this paper we are concerned with the challenging problem of producing a full image sequence of a deformable face given only an image and generic facial motions encoded by a set of sparse landmarks. To this end we build upon recent…

计算机视觉与模式识别 · 计算机科学 2019-04-29 Kritaphat Songsri-in , Stefanos Zafeiriou

COVID-19 has made video communication one of the most important modes of information exchange. While extensive research has been conducted on the optimization of the video streaming pipeline, in particular the development of novel video…

图像与视频处理 · 电气工程与系统科学 2021-01-11 Roshan Prabhakar , Shubham Chandak , Carina Chiu , Renee Liang , Huong Nguyen , Kedar Tatwawadi , Tsachy Weissman

Predominant techniques on talking head generation largely depend on 2D information, including facial appearances and motions from input face images. Nevertheless, dense 3D facial geometry, such as pixel-wise depth, plays a critical role in…

计算机视觉与模式识别 · 计算机科学 2023-12-12 Fa-Ting Hong , Li Shen , Dan Xu

We propose a novel landmarks-assisted collaborative end-to-end deep framework for automatic 4D FER. Using 4D face scan data, we calculate its various geometrical images, and afterwards use rank pooling to generate their dynamic images…

计算机视觉与模式识别 · 计算机科学 2020-02-10 Muzammil Behzad , Nhat Vo , Xiaobai Li , Guoying Zhao

Real-time 3D scene reconstruction from RGB-D sensor data, as well as the exploration of such data in VR/AR settings, has seen tremendous progress in recent years. The combination of both these components into telepresence systems, however,…

人机交互 · 计算机科学 2019-04-15 Patrick Stotko , Stefan Krumpen , Matthias B. Hullin , Michael Weinmann , Reinhard Klein

Facial landmarks refer to the localization of fundamental facial points on face images. There have been a tremendous amount of attempts to detect these points from facial images however, there has never been an attempt to synthesize a…

图像与视频处理 · 电气工程与系统科学 2018-02-02 Shabab Bazrafkan , Hossein Javidnia , Peter Corcoran