English
Related papers

Related papers: VFHQ: A High-Quality Dataset and Benchmark for Vid…

200 papers

Existing face super-resolution (FSR) methods have made significant advancements, but they primarily super-resolve face with limited visual information, original pixel-wise space in particular, commonly overlooking the pluralistic clues,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-15 Chenyang Wang , Wenjie An , Kui Jiang , Xianming Liu , Junjun Jiang

Finding the right sound effects (SFX) to match moments in a video is a difficult and time-consuming task, and relies heavily on the quality and completeness of text metadata. Retrieving high-quality (HQ) SFX using a video frame directly as…

Sound · Computer Science 2023-08-21 Julia Wilkins , Justin Salamon , Magdalena Fuentes , Juan Pablo Bello , Oriol Nieto

While recent works on blind face image restoration have successfully produced impressive high-quality (HQ) images with abundant details from low-quality (LQ) input images, the generated content may not accurately reflect the real appearance…

Computer Vision and Pattern Recognition · Computer Science 2024-12-09 Chi-Wei Hsiao , Yu-Lun Liu , Cheng-Kun Yang , Sheng-Po Kuo , Kevin Jou , Chia-Ping Chen

In recent years, speaker recognition systems based on raw waveform inputs have received increasing attention. However, the performance of such systems are typically inferior to the state-of-the-art handcrafted feature-based counterparts,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-03-30 Jee-weon Jung , You Jin Kim , Hee-Soo Heo , Bong-Jin Lee , Youngki Kwon , Joon Son Chung

With recent advances in computer vision and graphics, it is now possible to generate videos with extremely realistic synthetic faces, even in real time. Countless applications are possible, some of which raise a legitimate alarm, calling…

Computer Vision and Pattern Recognition · Computer Science 2018-03-28 Andreas Rössler , Davide Cozzolino , Luisa Verdoliva , Christian Riess , Justus Thies , Matthias Nießner

Recent video super-resolution (VSR) approaches use deep neural networks to enhance low-quality input videos and recover visual detail, with diffusion-based methods in particular showing promising results. In this paper, we investigate…

Image and Video Processing · Electrical Eng. & Systems 2026-05-26 Benjamin Herb , Steve Göring , Alexander Raake , Rakesh Rao Ramachandra Rao

We propose a novel method to use both audio and a low-resolution image to perform extreme face super-resolution (a 16x increase of the input size). When the resolution of the input image is very low (e.g., 8x8 pixels), the loss of…

Computer Vision and Pattern Recognition · Computer Science 2020-04-03 Givi Meishvili , Simon Jenni , Paolo Favaro

High-quality human reconstruction and photo-realistic rendering of a dynamic scene is a long-standing problem in computer vision and graphics. Despite considerable efforts invested in developing various capture systems and reconstruction…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Xiaoyun Zheng , Liwei Liao , Xufeng Li , Jianbo Jiao , Rongjie Wang , Feng Gao , Shiqi Wang , Ronggang Wang

Video Super-Resolution (VSR) aims to restore high-resolution (HR) videos from low-resolution (LR) videos. Existing VSR techniques usually recover HR frames by extracting pertinent textures from nearby frames with known degradation…

Image and Video Processing · Electrical Eng. & Systems 2023-01-02 Zhongwei Qiu , Huan Yang , Jianlong Fu , Daochang Liu , Chang Xu , Dongmei Fu

Video generation has witnessed significant advancements, yet evaluating these models remains a challenge. A comprehensive evaluation benchmark for video generation is indispensable for two reasons: 1) Existing metrics do not fully align…

Computer Vision and Pattern Recognition · Computer Science 2023-12-01 Ziqi Huang , Yinan He , Jiashuo Yu , Fan Zhang , Chenyang Si , Yuming Jiang , Yuanhan Zhang , Tianxing Wu , Qingyang Jin , Nattapol Chanpaisit , Yaohui Wang , Xinyuan Chen , Limin Wang , Dahua Lin , Yu Qiao , Ziwei Liu

High dynamic range (HDR) video reconstruction from sequences captured with alternating exposures is a very challenging problem. Existing methods often align low dynamic range (LDR) input sequence in the image space using optical flow, and…

Computer Vision and Pattern Recognition · Computer Science 2021-08-24 Guanying Chen , Chaofeng Chen , Shi Guo , Zhetong Liang , Kwan-Yee K. Wong , Lei Zhang

Blind face restoration (BFR) on images has significantly progressed over the last several years, while real-world video face restoration (VFR), which is more challenging for more complex face motions such as moving gaze directions and…

Computer Vision and Pattern Recognition · Computer Science 2024-05-07 Ziyan Chen , Jingwen He , Xinqi Lin , Yu Qiao , Chao Dong

Vision-language models (VLMs) excel in various visual benchmarks but are often constrained by the lack of high-quality visual fine-tuning data. To address this challenge, we introduce VisCon-100K, a novel dataset derived from interleaved…

Computation and Language · Computer Science 2025-02-25 Gokul Karthik Kumar , Iheb Chaabane , Kebin Wu

The volumetric representation of human interactions is one of the fundamental domains in the development of immersive media productions and telecommunication applications. Particularly in the context of the rapid advancement of Extended…

Computer Vision and Pattern Recognition · Computer Science 2024-02-15 Fatemeh Ghorbani Lohesara , Davi Rabbouni Freitas , Christine Guillemot , Karen Eguiazarian , Sebastian Knorr

Compressed video super-resolution (SR) aims to generate high-resolution (HR) videos from the corresponding low-resolution (LR) compressed videos. Recently, some compressed video SR methods attempt to exploit the spatio-temporal information…

Computer Vision and Pattern Recognition · Computer Science 2026-04-23 Qiang Zhu , Fan Zhang , Feiyu Chen , Shuyuan Zhu , David Bull , Bing Zeng

General-purpose audio representations aim to map acoustically variable instances of the same event to nearby points, resolving content identity in a zero-shot setting. Unlike supervised classification benchmarks that measure adaptability…

Sound · Computer Science 2025-12-12 Maris Basha , Anja Zai , Sabine Stoll , Richard Hahnloser

Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Hanyue Lou , Jinxiu Liang , Minggui Teng , Yi Wang , Boxin Shi

In this paper, we consider the problem of reference-based video super-resolution(RefVSR), i.e., how to utilize a high-resolution (HR) reference frame to super-resolve a low-resolution (LR) video sequence. The existing approaches to RefVSR…

Computer Vision and Pattern Recognition · Computer Science 2021-10-29 Yaping Zhao , Mengqi Ji , Ruqi Huang , Bin Wang , Shengjin Wang

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

Generating natural language questions from visual scenes, known as Visual Question Generation (VQG), has been explored in the recent past where large amounts of meticulously labeled data provide the training corpus. However, in practice, it…

Computer Vision and Pattern Recognition · Computer Science 2023-01-09 Anurag Roy , David Johnson Ekka , Saptarshi Ghosh , Abir Das