English
Related papers

Related papers: VFHQ: A High-Quality Dataset and Benchmark for Vid…

200 papers

The threat of Audio-Video (AV) forgery is rapidly evolving beyond human-centric deepfakes to include more diverse manipulations across complex natural scenes. However, existing benchmarks are still confined to DeepFake-based forgeries and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Shuhan Xia , Peipei Li , Xuannan Liu , Dongsen Zhang , Xinyu Guo , Zekun Li

Large vision-language models (LVLMs) have demonstrated remarkable achievements, yet the generation of non-factual responses remains prevalent in fact-seeking question answering (QA). Current multimodal fact-seeking benchmarks primarily…

Computation and Language · Computer Science 2025-03-11 Yanling Wang , Yihan Zhao , Xiaodong Chen , Shasha Guo , Lixin Liu , Haoyang Li , Yong Xiao , Jing Zhang , Qi Li , Ke Xu

Recently reported state-of-the-art results in visual speech recognition (VSR) often rely on increasingly large amounts of video data, while the publicly available transcribed video datasets are limited in size. In this paper, for the first…

Super-resolution algorithms often struggle with images from surveillance environments due to adverse conditions such as unknown degradation, variations in pose, irregular illumination, and occlusions. However, acquiring multiple images,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Marcelo dos Santos , Rayson Laroca , Rafael O. Ribeiro , João C. Neves , David Menotti

In face detection, low-resolution faces, such as numerous small faces of a human group in a crowded scene, are common in dense face prediction tasks. They usually contain limited visual clues and make small faces less distinguishable from…

Computer Vision and Pattern Recognition · Computer Science 2023-06-06 Guangtao Wang , Jun Li , Jie Xie , Jianhua Xu , Bo Yang

The VoxCeleb Speaker Recognition Challenge (VoxSRC) at Interspeech 2020 offers a challenging evaluation for speaker recognition systems, which includes celebrities playing different parts in movies. The goal of this work is robust speaker…

Sound · Computer Science 2020-10-30 Yoohwan Kwon , Hee-Soo Heo , Bong-Jin Lee , Joon Son Chung

High frame rates have been known to enhance the perceived visual quality of specific video content. However, the lack of investigation of high frame rates has restricted the expansion of this research field particularly in the context of…

Image and Video Processing · Electrical Eng. & Systems 2020-06-05 Tariq Rahim , Muhammad Arslan Usman , Soo Young Shin

Speaker-independent VSR is a complex task that involves identifying spoken words or phrases from video recordings of a speaker's facial movements. Over the years, there has been a considerable amount of research in the field of VSR…

Computer Vision and Pattern Recognition · Computer Science 2023-12-13 Praneeth Nemani , G. Sai Krishna , Supriya Kundrapu

With the rapid advancement of technologies like text-to-speech (TTS) and voice conversion (VC), detecting deepfake voices has become increasingly crucial. However, both academia and industry lack a comprehensive and intuitive benchmark for…

Sound · Computer Science 2024-09-11 Ziwei Yan , Yanjie Zhao , Haoyu Wang

Current benchmarks for facial expression recognition (FER) mainly focus on static images, while there are limited datasets for FER in videos. It is still ambiguous to evaluate whether performances of existing methods remain satisfactory in…

Computer Vision and Pattern Recognition · Computer Science 2022-03-22 Yan Wang , Yixuan Sun , Yiwen Huang , Zhongying Liu , Shuyong Gao , Wei Zhang , Weifeng Ge , Wenqiang Zhang

Talking face generation (TFG) allows for producing lifelike talking videos of any character using only facial images and accompanying text. Abuse of this technology could pose significant risks to society, creating the urgent need for…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Xiaocan Chen , Qilin Yin , Jiarui Liu , Wei Lu , Xiangyang Luo , Jiantao Zhou

Multimodal manipulations (also known as audio-visual deepfakes) make it difficult for unimodal deepfake detectors to detect forgeries in multimedia content. To avoid the spread of false propaganda and fake news, timely detection is crucial.…

Computer Vision and Pattern Recognition · Computer Science 2025-11-24 Sahibzada Adil Shahzad , Ammarah Hashmi , Yan-Tsung Peng , Yu Tsao , Hsin-Min Wang

Visual speech recognition (VSR) aims to recognize the content of speech based on lip movements, without relying on the audio stream. Advances in deep learning and the availability of large audio-visual datasets have led to the development…

Computer Vision and Pattern Recognition · Computer Science 2022-11-01 Pingchuan Ma , Stavros Petridis , Maja Pantic

Deepfakes are realistic face manipulations that can pose serious threats to security, privacy, and trust. Existing methods mostly treat this task as binary classification, which uses digital labels or mask signals to train the detection…

Computer Vision and Pattern Recognition · Computer Science 2024-02-08 Ke Sun , Shen Chen , Taiping Yao , Haozhe Yang , Xiaoshuai Sun , Shouhong Ding , Rongrong Ji

The size of training dataset is known to be among the most dominating aspects of training high-performance face recognition embedding model. Building a large dataset from scratch could be cumbersome and time-intensive, while combining…

Computer Vision and Pattern Recognition · Computer Science 2023-05-25 Chiyoung Song , Dongjae Lee

Visual Question Answering (VQA) is a challenging task that requires the joint understanding of natural language and visual content. While early research primarily focused on recognizing objects and scene context, it often overlooked scene…

Recently, high-quality video conferencing with fewer transmission bits has become a very hot and challenging problem. We propose FAIVConf, a specially designed video compression framework for video conferencing, based on the effective…

Image and Video Processing · Electrical Eng. & Systems 2022-07-12 Zhengang Li , Sheng Lin , Shan Liu , Songnan Li , Xue Lin , Wei Wang , Wei Jiang

The proliferation of deepfake faces poses huge potential negative impacts on our daily lives. Despite substantial advancements in deepfake detection over these years, the generalizability of existing methods against forgeries from unseen…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Kaiqing Lin , Yuzhen Lin , Weixiang Li , Taiping Yao , Bin Li

Representing human performance at high-fidelity is an essential building block in diverse applications, such as film production, computer games or videoconferencing. To close the gap to production-level quality, we introduce HumanRF, a 4D…

Computer Vision and Pattern Recognition · Computer Science 2023-05-12 Mustafa Işık , Martin Rünz , Markos Georgopoulos , Taras Khakhulin , Jonathan Starck , Lourdes Agapito , Matthias Nießner

The rapid surge of text-to-speech and face-voice reenactment models makes video fabrication easier and highly realistic. To encounter this problem, we require datasets that rich in type of generation methods and perturbation strategy which…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Zhixi Cai , Kartik Kuckreja , Shreya Ghosh , Akanksha Chuchra , Muhammad Haris Khan , Usman Tariq , Tom Gedeon , Abhinav Dhall
‹ Prev 1 8 9 10 Next ›