中文
相关论文

相关论文: Learning to Have an Ear for Face Super-Resolution

200 篇论文

In this paper, we present multimodal deep neural network frameworks for age and gender classification, which take input a profile face image as well as an ear image. Our main objective is to enhance the accuracy of soft biometric trait…

计算机视觉与模式识别 · 计算机科学 2019-07-25 Dogucan Yaman , Fevziye Irem Eyiokur , Hazım Kemal Ekenel

Enhancing speech quality is an indispensable yet difficult task as it is often complicated by a range of degradation factors. In addition to additive noise, reverberation, clipping, and speech attenuation can all adversely affect speech…

音频与语音处理 · 电气工程与系统科学 2023-06-01 Jaeuk Byun , Youna Ji , Soo Whan Chung , Soyeon Choe , Min Seok Choi

We present a practical approach to capturing ear-to-ear face models comprising both 3D meshes and intrinsic textures (i.e. diffuse and specular albedo). Our approach is a hybrid of geometric and photometric methods and requires no geometric…

计算机视觉与模式识别 · 计算机科学 2016-09-09 Alassane Seck , William A. P. Smith , Arnaud Dessein , Bernard Tiddeman , Hannah Dee , Abhishek Dutta

Speech enhancement and speech separation are two related tasks, whose purpose is to extract either one or more target speech signals, respectively, from a mixture of sounds generated by several sources. Traditionally, these tasks have been…

音频与语音处理 · 电气工程与系统科学 2021-03-16 Daniel Michelsanti , Zheng-Hua Tan , Shi-Xiong Zhang , Yong Xu , Meng Yu , Dong Yu , Jesper Jensen

Recently, 3D face reconstruction from a single image has achieved great success with the help of deep learning and shape prior knowledge, but they often fail to produce accurate geometry details. On the other hand, photometric stereo…

计算机视觉与模式识别 · 计算机科学 2020-03-30 Xueying Wang , Yudong Guo , Bailin Deng , Juyong Zhang

Audio-visual automatic speech recognition is a promising approach to robust ASR under noisy conditions. However, up until recently it had been traditionally studied in isolation assuming the video of a single speaking face matches the…

音频与语音处理 · 电气工程与系统科学 2022-05-13 Otavio Braga , Olivier Siohan

Super-resolution algorithms often struggle with images from surveillance environments due to adverse conditions such as unknown degradation, variations in pose, irregular illumination, and occlusions. However, acquiring multiple images,…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Marcelo dos Santos , Rayson Laroca , Rafael O. Ribeiro , João C. Neves , David Menotti

The lack of resolution has a negative impact on the performance of image-based biometrics. Many applications which are becoming ubiquitous in mobile devices do not operate in a controlled environment, and their performance significantly…

计算机视觉与模式识别 · 计算机科学 2022-04-13 Fernando Alonso-Fernandez , Reuben A. Farrugia , Julian Fierrez , Josef Bigun

Speaker profiling, which aims to estimate speaker characteristics such as age and height, has a wide range of applications inforensics, recommendation systems, etc. In this work, we propose a semisupervised learning approach to mitigate the…

音频与语音处理 · 电气工程与系统科学 2021-10-27 Shangeth Rajaa , Pham Van Tung , Chng Eng Siong

Low-resolution face recognition (LRFR) has received increasing attention over the past few years. Its applications lie widely in the real-world environment when high-resolution or high-quality images are hard to capture. One of the biggest…

计算机视觉与模式识别 · 计算机科学 2019-04-01 Pei Li , Loreto Prieto , Domingo Mery , Patrick Flynn

Existing face relighting methods often struggle with two problems: maintaining the local facial details of the subject and accurately removing and synthesizing shadows in the relit image, especially hard shadows. We propose a novel deep…

计算机视觉与模式识别 · 计算机科学 2021-06-08 Andrew Hou , Ze Zhang , Michel Sarkis , Ning Bi , Yiying Tong , Xiaoming Liu

Novel text-to-speech systems can generate entirely new voices that were not seen during training. However, it remains a difficult task to efficiently create personalized voices from a high-dimensional speaker space. In this work, we use…

We evaluate the information that can unintentionally leak into the low dimensional output of a neural network, by reconstructing an input image from a 40- or 32-element feature vector that intends to only describe abstract attributes of a…

计算机视觉与模式识别 · 计算机科学 2025-03-13 Kathleen Anderson , Thomas Martinetz

Today, Multi-View Stereo techniques are able to reconstruct robust and detailed 3D models, especially when starting from high-resolution images. However, there are cases in which the resolution of input images is relatively low, for…

计算机视觉与模式识别 · 计算机科学 2021-07-29 Eugenio Lomurno , Andrea Romanoni , Matteo Matteucci

We introduce the first method for audio-driven universal photorealistic avatar synthesis, combining a person-agnostic speech model with our novel Universal Head Avatar Prior (UHAP). UHAP is trained on cross-identity multi-view videos. In…

计算机视觉与模式识别 · 计算机科学 2025-09-24 Kartik Teotia , Helge Rhodin , Mohit Mendiratta , Hyeongwoo Kim , Marc Habermann , Christian Theobalt

Audio captioning aims to generate text descriptions of audio clips. In the real world, many objects produce similar sounds. How to accurately recognize ambiguous sounds is a major challenge for audio captioning. In this work, inspired by…

音频与语音处理 · 电气工程与系统科学 2023-05-30 Xubo Liu , Qiushi Huang , Xinhao Mei , Haohe Liu , Qiuqiang Kong , Jianyuan Sun , Shengchen Li , Tom Ko , Yu Zhang , Lilian H. Tang , Mark D. Plumbley , Volkan Kılıç , Wenwu Wang

The objective of this paper is to separate a target speaker's speech from a mixture of two speakers using a deep audio-visual speech separation network. Unlike previous works that used lip movement on video clips or pre-enrolled speaker…

声音 · 计算机科学 2020-11-05 Soo-Whan Chung , Soyeon Choe , Joon Son Chung , Hong-Goo Kang

We present a framework for generating full-bodied photorealistic avatars that gesture according to the conversational dynamics of a dyadic interaction. Given speech audio, we output multiple possibilities of gestural motion for an…

计算机视觉与模式识别 · 计算机科学 2024-01-04 Evonne Ng , Javier Romero , Timur Bagautdinov , Shaojie Bai , Trevor Darrell , Angjoo Kanazawa , Alexander Richard

We present a method that generates expressive talking heads from a single facial image with audio as the only input. In contrast to previous approaches that attempt to learn direct mappings from audio to raw pixels or points for creating…

计算机视觉与模式识别 · 计算机科学 2021-02-26 Yang Zhou , Xintong Han , Eli Shechtman , Jose Echevarria , Evangelos Kalogerakis , Dingzeyu Li

Face Super-Resolution (SR) is a subfield of the SR domain that specifically targets the reconstruction of face images. The main challenge of face SR is to restore essential facial features without distortion. We propose a novel face SR…

计算机视觉与模式识别 · 计算机科学 2019-08-23 Deokyun Kim , Minseon Kim , Gihyun Kwon , Dae-Shik Kim