English
Related papers

Related papers: DH-FaceVid-1K: A Large-Scale High-Quality Dataset …

200 papers

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Haopeng Fang , Di Qiu , Binjie Mao , He Tang

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang

Deep learning for human action recognition in videos is making significant progress, but is slowed down by its dependency on expensive manual labeling of large video collections. In this work, we investigate the generation of synthetic…

Computer Vision and Pattern Recognition · Computer Science 2017-07-20 César Roberto de Souza , Adrien Gaidon , Yohann Cabon , Antonio Manuel López Peña

To mitigate the risk of harmful outputs from large vision models (LVMs), we introduce the SafeSora dataset to promote research on aligning text-to-video generation with human values. This dataset encompasses human preferences in…

Computer Vision and Pattern Recognition · Computer Science 2024-06-21 Josef Dai , Tianle Chen , Xuyao Wang , Ziran Yang , Taiye Chen , Jiaming Ji , Yaodong Yang

Creating and labelling datasets of videos for use in training Human Activity Recognition models is an arduous task. In this paper, we approach this by using 3D rendering tools to generate a synthetic dataset of videos, and show that a…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Ollie Matthews , Koki Ryu , Tarun Srivastava

In this paper, we present our approach to the DataCV ICCV Challenge, which centers on building a high-quality face dataset to train a face recognition model. The constructed dataset must not contain identities overlapping with any existing…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Feiran Li , Qianqian Xu , Shilong Bao , Boyu Han , Zhiyong Yang , Qingming Huang

Generating synthetic datasets for training face recognition models is challenging because dataset generation entails more than creating high fidelity images. It involves generating multiple images of same subjects under different factors…

Computer Vision and Pattern Recognition · Computer Science 2023-04-17 Minchul Kim , Feng Liu , Anil Jain , Xiaoming Liu

Generating long, high-quality videos remains a challenge due to the complex interplay of spatial and temporal dynamics and hardware limitations. In this work, we introduce MaskFlow, a unified video generation framework that combines…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Michael Fuest , Vincent Tao Hu , Björn Ommer

The recent and increasing interest in video-language research has driven the development of large-scale datasets that enable data-intensive machine learning techniques. In comparison, limited effort has been made at assessing the fitness of…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Mattia Soldan , Alejandro Pardo , Juan León Alcázar , Fabian Caba Heilbron , Chen Zhao , Silvio Giancola , Bernard Ghanem

In this era, the success of large language models and text-to-image models can be attributed to the driving force of large-scale datasets. However, in the realm of 3D vision, while significant progress has been achieved in object-centric…

Computer Vision and Pattern Recognition · Computer Science 2025-05-06 Chenghong Li , Hongjie Liao , Yihao Zhi , Xihe Yang , Zhengwentai Sun , Jiahao Chang , Shuguang Cui , Xiaoguang Han

In this paper, we present a large-scale detailed 3D face dataset, FaceScape, and the corresponding benchmark to evaluate single-view facial 3D reconstruction. By training on FaceScape data, a novel algorithm is proposed to predict elaborate…

Computer Vision and Pattern Recognition · Computer Science 2023-09-19 Hao Zhu , Haotian Yang , Longwei Guo , Yidi Zhang , Yanru Wang , Mingkai Huang , Menghua Wu , Qiu Shen , Ruigang Yang , Xun Cao

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

Computer Vision and Pattern Recognition · Computer Science 2024-06-13 Jiadong Liang , Feng Lu

Deep networks trained on millions of facial images are believed to be closely approaching human-level performance in face recognition. However, open world face recognition still remains a challenge. Although, 3D face recognition has an…

Computer Vision and Pattern Recognition · Computer Science 2020-12-03 Syed Zulqarnain Gilani , Ajmal Mian

Short-video platforms show an increasing impact on people's daily lives nowadays, with billions of active users spending plenty of time each day. The interactions between users and online platforms give rise to many scientific problems…

Multimedia · Computer Science 2025-02-11 Yu Shang , Chen Gao , Nian Li , Yong Li

Talking head video generation aims to generate a realistic talking head video that preserves the person's identity from a source image and the motion from a driving video. Despite the promising progress made in the field, it remains a…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Shuling Zhao , Fa-Ting Hong , Xiaoshui Huang , Dan Xu

The rise of deep generative models has greatly advanced video compression, reshaping the paradigm of face video coding through their powerful capability for semantic-aware representation and lifelike synthesis. Generative Face Video Coding…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Bolin Chen , Shanzhi Yin , Goluck Konuko , Giuseppe Valenzise , Zihan Zhang , Shiqi Wang , Yan Ye

Modeling the reactive tempo of human conversation remains difficult because most audio-visual datasets portray isolated speakers delivering short monologues. We introduce \textbf{Face-to-Face with Jimmy Fallon (F2F-JF)}, a 70-hour, 14k-clip…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Ernie Chu , Vishal M. Patel

Predominant techniques on talking head generation largely depend on 2D information, including facial appearances and motions from input face images. Nevertheless, dense 3D facial geometry, such as pixel-wise depth, plays a critical role in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Fa-Ting Hong , Li Shen , Dan Xu

Human generation has achieved significant progress. Nonetheless, existing methods still struggle to synthesize specific regions such as faces and hands. We argue that the main reason is rooted in the training data. A holistic human dataset…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Jianglin Fu , Shikai Li , Yuming Jiang , Kwan-Yee Lin , Wayne Wu , Ziwei Liu

This paper studies how to synthesize face images of non-existent persons, to create a dataset that allows effective training of face recognition (FR) models. Besides generating realistic face images, two other important goals are: 1) the…

Computer Vision and Pattern Recognition · Computer Science 2025-02-10 Haiyu Wu , Jaskirat Singh , Sicong Tian , Liang Zheng , Kevin W. Bowyer