English
Related papers

Related papers: Learning to Have an Ear for Face Super-Resolution

200 papers

Implicit neural representations (INRs) are a rapidly growing research field, which provides alternative ways to represent multimedia signals. Recent applications of INRs include image super-resolution, compression of high-dimensional…

In this paper, we explore the correlation between different visual biometric modalities. For this purpose, we present an end-to-end deep neural network model that learns a mapping between the biometric modalities. Namely, our goal is to…

Computer Vision and Pattern Recognition · Computer Science 2020-06-04 Dogucan Yaman , Fevziye Irem Eyiokur , Hazım Kemal Ekenel

With the advance in self-supervised learning for audio and visual modalities, it has become possible to learn a robust audio-visual speech representation. This would be beneficial for improving the audio-visual speech recognition (AVSR)…

Image and Video Processing · Electrical Eng. & Systems 2022-07-12 Zi-Qiang Zhang , Jie Zhang , Jian-Shu Zhang , Ming-Hui Wu , Xin Fang , Li-Rong Dai

In this study, we propose AniPortrait, a novel framework for generating high-quality animation driven by audio and a reference portrait image. Our methodology is divided into two stages. Initially, we extract 3D intermediate representations…

Computer Vision and Pattern Recognition · Computer Science 2024-03-27 Huawei Wei , Zejun Yang , Zhisheng Wang

Preserving face identity is a critical yet persistent challenge in diffusion-based image restoration. While reference faces offer a path forward, existing reference-based methods often fail to fully exploit their potential. This paper…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Mo Zhou , Keren Ye , Viraj Shah , Kangfu Mei , Mauricio Delbracio , Peyman Milanfar , Vishal M. Patel , Hossein Talebi

Nowadays, due to the ubiquitous visual media there are vast amounts of already available high-resolution (HR) face images. Therefore, for super-resolving a given very low-resolution (LR) face image of a person it is very likely to find…

Computer Vision and Pattern Recognition · Computer Science 2019-06-18 Berk Dogan , Shuhang Gu , Radu Timofte

Speechreading or lipreading is the technique of understanding and getting phonetic features from a speaker's visual features such as movement of lips, face, teeth and tongue. It has a wide range of multimedia applications such as in…

In this paper, we present a detailed analysis on extracting soft biometric traits, age and gender, from ear images. Although there have been a few previous work on gender classification using ear images, to the best of our knowledge, this…

Computer Vision and Pattern Recognition · Computer Science 2018-06-18 Dogucan Yaman , Fevziye Irem Eyiokur , Nurdan Sezgin , Hazım Kemal Ekenel

Recent advances have demonstrated compelling capabilities in synthesizing real individuals into generated videos, reflecting the growing demand for identity-aware content creation. Nevertheless, an openly accessible framework enabling…

Computer Vision and Pattern Recognition · Computer Science 2026-03-26 Yingjie Chen , Shilun Lin , Cai Xing , Binxin Yang , Long Zhou , Qixin Yan , Wenjing Wang , Dingming Liu , Hao Liu , Chen Li , Jing Lyu

The intuitive interaction between the audio and visual modalities is valuable for cross-modal self-supervised learning. This concept has been demonstrated for generic audiovisual tasks like video action recognition and acoustic scene…

Audio and Speech Processing · Electrical Eng. & Systems 2020-07-14 Abhinav Shukla , Stavros Petridis , Maja Pantic

Ultrasound tongue imaging is widely used for speech production research, and it has attracted increasing attention as its potential applications seem to be evident in many different fields, such as the visual biofeedback tool for second…

Computer Vision and Pattern Recognition · Computer Science 2021-01-28 Kele Xu , Tamas Gábor Csapó , Ming Feng

In medical imaging, manual annotations can be expensive to acquire and sometimes infeasible to access, making conventional deep learning-based models difficult to scale. As a result, it would be beneficial if useful representations could be…

Computer Vision and Pattern Recognition · Computer Science 2020-08-18 Jianbo Jiao , Yifan Cai , Mohammad Alsharid , Lior Drukker , Aris T. Papageorghiou , J. Alison Noble

Training audio-to-image generative models requires an abundance of diverse audio-visual pairs that are semantically aligned. Such data is almost always curated from in-the-wild videos, given the cross-modal semantic correspondence that is…

Sound · Computer Science 2025-01-10 Darius Petermann , Mahdi M. Kalayeh

Speech representation and modelling in high-dimensional spaces of acoustic waveforms, or a linear transformation thereof, is investigated with the aim of improving the robustness of automatic speech recognition to additive noise. The…

Computation and Language · Computer Science 2015-03-31 Matthew Ager , Zoran Cvetkovic , Peter Sollich

Nowadays, we have witnessed the early progress on learning the association between voice and face automatically, which brings a new wave of studies to the computer vision community. However, most of the prior arts along this line (a) merely…

Computer Vision and Pattern Recognition · Computer Science 2021-03-15 Peisong Wen , Qianqian Xu , Yangbangyan Jiang , Zhiyong Yang , Yuan He , Qingming Huang

The light stage has been widely used in computer graphics for the past two decades, primarily to enable the relighting of human faces. By capturing the appearance of the human subject under different light sources, one obtains the light…

We consider the task of upscaling a low resolution thumbnail image of a person, to a higher resolution image, which preserves the person's identity and other attributes. Since the thumbnail image is of low resolution, many higher resolution…

Computer Vision and Pattern Recognition · Computer Science 2021-06-01 Noam Gat , Sagie Benaim , Lior Wolf

We present a simple method to reconstruct a high-resolution video from a face-video, where the identity of a person is obscured by pixelization. This concealment method is popular because the viewer can still perceive a human face figure…

Computer Vision and Pattern Recognition · Computer Science 2020-09-30 Maayan Shuvi , Noa Fish , Kfir Aberman , Ariel Shamir , Daniel Cohen-Or

We present a neural network-based simulation super-resolution framework that can efficiently and realistically enhance a facial performance produced by a low-cost, realtime physics-based simulation to a level of detail that closely…

We present a novel approach for super-resolution that utilizes implicit neural representation (INR) to effectively reconstruct and enhance low-resolution videos and images. By leveraging the capacity of neural networks to implicitly encode…

Computer Vision and Pattern Recognition · Computer Science 2025-03-07 Mary Aiyetigbo , Wanqi Yuan , Feng Luo , Nianyi Li
‹ Prev 1 3 4 5 6 7 10 Next ›