中文
相关论文

相关论文: VFHQ: A High-Quality Dataset and Benchmark for Vid…

200 篇论文

In this paper, we provide a large audio-visual speaker recognition dataset, VoxBlink2, which includes approximately 10M utterances with videos from 110K+ speakers in the wild. This dataset represents a significant expansion over the…

音频与语音处理 · 电气工程与系统科学 2024-07-17 Yuke Lin , Ming Cheng , Fulin Zhang , Yingying Gao , Shilei Zhang , Ming Li

Talking-head videos constitute a predominant content type in real-time communication, yet publicly available datasets for video processing research in this domain remain scarce and limited in signal fidelity. In this paper, we open-source a…

计算机视觉与模式识别 · 计算机科学 2026-03-31 Babak Naderi , Ross Cutler

We present VQTalker, a Vector Quantization-based framework for multilingual talking head generation that addresses the challenges of lip synchronization and natural motion across diverse languages. Our approach is grounded in the phonetic…

计算机视觉与模式识别 · 计算机科学 2024-12-19 Tao Liu , Ziyang Ma , Qi Chen , Feilong Chen , Shuai Fan , Xie Chen , Kai Yu

Face super-resolution (FSR), also known as face hallucination, which is aimed at enhancing the resolution of low-resolution (LR) face images to generate high-resolution (HR) face images, is a domain-specific image super-resolution problem.…

计算机视觉与模式识别 · 计算机科学 2021-09-02 Junjun Jiang , Chenyang Wang , Xianming Liu , Jiayi Ma

Detecting partial deepfake speech is challenging because manipulations occur only in short regions while the surrounding audio remains authentic. However, existing detection methods are fundamentally limited by the quality of available…

声音 · 计算机科学 2025-12-16 Menglu Li , Majd Alber , Ramtin Asgarianamiri , Lian Zhao , Xiao-Ping Zhang

Video super-resolution (VSR) aims to enhance low-resolution videos by leveraging both spatial and temporal information. While deep learning has led to impressive progress, it typically requires centralized data, which raises privacy…

计算机视觉与模式识别 · 计算机科学 2026-02-04 Ali Mollaahmadi Dehaghi , Hossein KhademSohi , Reza Razavi , Steve Drew , Mohammad Moshirpour

Robots are becoming everyday devices, increasing their interaction with humans. To make human-machine interaction more natural, cognitive features like Visual Voice Activity Detection (VVAD), which can detect whether a person is speaking or…

计算机视觉与模式识别 · 计算机科学 2024-01-03 Adrian Lubitz , Matias Valdenegro-Toro , Frank Kirchner

Existing automatic captioning methods for visual content face challenges such as lack of detail, content hallucination, and poor instruction following. In this work, we propose VisualFactChecker (VFC), a flexible training-free pipeline that…

计算机视觉与模式识别 · 计算机科学 2024-05-01 Yunhao Ge , Xiaohui Zeng , Jacob Samuel Huffman , Tsung-Yi Lin , Ming-Yu Liu , Yin Cui

In recent years, heatmap regression based models have shown their effectiveness in face alignment and pose estimation. However, Conventional Heatmap Regression (CHR) is not accurate nor stable when dealing with high-resolution facial…

计算机视觉与模式识别 · 计算机科学 2018-11-26 Ying Tai , Yicong Liang , Xiaoming Liu , Lei Duan , Jilin Li , Chengjie Wang , Feiyue Huang , Yu Chen

This paper explores sentence-level multilingual Visual Speech Recognition (VSR) that can recognize different languages with a single trained model. As the massive multilingual modeling of visual data requires huge computational costs, we…

音频与语音处理 · 电气工程与系统科学 2024-07-19 Minsu Kim , Jeong Hun Yeo , Se Jin Park , Hyeongseop Rha , Yong Man Ro

Velopharyngeal dysfunction (VPD) is characterized by inadequate velopharyngeal closure during speech and often causes hypernasality and reduced intelligibility. Although speech-based machine learning models can perform well under…

音频与语音处理 · 电气工程与系统科学 2026-03-19 Weixin Liu , Bowen Qu , Amy Stone , Maria E. Powell , Shama Dufresne , Stephane Braun , Izabela Galdyn , Michael Golinko , Bradley Malin , Zhijun Yin , Matthew E. Pontell

Human head detection, keypoint estimation, and 3D head model fitting are essential tasks with many applications. However, traditional real-world datasets often suffer from bias, privacy, and ethical concerns, and they have been recorded in…

计算机视觉与模式识别 · 计算机科学 2024-12-06 Orest Kupyn , Eugene Khvedchenia , Christian Rupprecht

In this paper, we investigate the task of hallucinating an authentic high-resolution (HR) human face from multiple low-resolution (LR) video snapshots. We propose a pure transformer-based model, dubbed VidFace, to fully exploit the…

计算机视觉与模式识别 · 计算机科学 2021-06-01 Yuan Gan , Yawei Luo , Xin Yu , Bang Zhang , Yi Yang

Human pose and shape (HPS) estimation methods have been extensively studied, with many demonstrating high zero-shot performance on in-the-wild images and videos. However, these methods often struggle in challenging scenarios involving…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Yash Garg , Saketh Bachu , Arindam Dutta , Rohit Lal , Sarosij Bose , Calvin-Khang Ta , M. Salman Asif , Amit Roy-Chowdhury

Designing and manipulating virtual human heads is essential across various applications, including AR, VR, gaming, human-computer interaction and VFX. Traditional graphic-based approaches require manual effort and resources to achieve…

计算机视觉与模式识别 · 计算机科学 2024-07-02 Anirban Mukherjee , Venkat Suprabath Bitra , Vignesh Bondugula , Tarun Reddy Tallapureddy , Dinesh Babu Jayagopi

In recent years, deep learning has made great progress in many fields such as image recognition, natural language processing, speech recognition and video super-resolution. In this survey, we comprehensively investigate 33 state-of-the-art…

计算机视觉与模式识别 · 计算机科学 2022-03-17 Hongying Liu , Zhubo Ruan , Peng Zhao , Chao Dong , Fanhua Shang , Yuanyuan Liu , Linlin Yang , Radu Timofte

Recognizing human non-speech vocalizations is an important task and has broad applications such as automatic sound transcription and health condition monitoring. However, existing datasets have a relatively small number of vocal sound…

声音 · 计算机科学 2022-06-22 Yuan Gong , Jin Yu , James Glass

Video face re-aging deals with altering the apparent age of a person to the target age in videos. This problem is challenging due to the lack of paired video datasets maintaining temporal consistency in identity and age. Most re-aging…

计算机视觉与模式识别 · 计算机科学 2024-03-15 Abdul Muqeet , Kyuchul Lee , Bumsoo Kim , Yohan Hong , Hyungrae Lee , Woonggon Kim , KwangHee Lee

This paper studies how to synthesize face images of non-existent persons, to create a dataset that allows effective training of face recognition (FR) models. Besides generating realistic face images, two other important goals are: 1) the…

计算机视觉与模式识别 · 计算机科学 2025-02-10 Haiyu Wu , Jaskirat Singh , Sicong Tian , Liang Zheng , Kevin W. Bowyer

Deep learning-based face recognition continues to face challenges due to its reliance on huge datasets obtained from web crawling, which can be costly to gather and raise significant real-world privacy concerns. To address this issue, we…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Minsoo Kim , Min-Cheol Sagong , Gi Pyo Nam , Junghyun Cho , Ig-Jae Kim