English
Related papers

Related papers: High-Fidelity 3D Face Generation from Natural Lang…

200 papers

We present a method for fine-grained face manipulation. Given a face image with an arbitrary expression, our method can synthesize another arbitrary expression by the same person. This is achieved by first fitting a 3D face model and then…

Computer Vision and Pattern Recognition · Computer Science 2019-02-26 Zhenglin Geng , Chen Cao , Sergey Tulyakov

The goal of this work is to simultaneously generate natural talking faces and speech outputs from text. We achieve this by integrating Talking Face Generation (TFG) and Text-to-Speech (TTS) systems into a unified framework. We address the…

Computer Vision and Pattern Recognition · Computer Science 2024-05-17 Youngjoon Jang , Ji-Hoon Kim , Junseok Ahn , Doyeop Kwak , Hong-Sun Yang , Yoon-Cheol Ju , Il-Hwan Kim , Byeong-Yeol Kim , Joon Son Chung

We study the problem of creating high-fidelity and animatable 3D avatars from only textual descriptions. Existing text-to-avatar methods are either limited to static avatars which cannot be animated or struggle to generate animatable…

Graphics · Computer Science 2023-11-30 Jianfeng Zhang , Xuanmeng Zhang , Huichao Zhang , Jun Hao Liew , Chenxu Zhang , Yi Yang , Jiashi Feng

Impressive progress has been made in audio-driven 3D facial animation recently, but synthesizing 3D talking-head with rich emotion is still unsolved. This is due to the lack of 3D generative models and available 3D emotional dataset with…

Computer Vision and Pattern Recognition · Computer Science 2021-04-27 Qianyun Wang , Zhenfeng Fan , Shihong Xia

3D indoor scene generation is an important problem for the design of digital and real-world environments. To automate this process, a scene generation model should be able to not only generate plausible scene layouts, but also take into…

Computer Vision and Pattern Recognition · Computer Science 2025-03-26 Kelly O. Marshall , Omid Poursaeed , Sergiu Oprea , Amit Kumar , Anushrut Jignasu , Chinmay Hegde , Yilei Li , Rakesh Ranjan

Person-generic audio-driven face generation is a challenging task in computer vision. Previous methods have achieved remarkable progress in audio-visual synchronization, but there is still a significant gap between current results and…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Xiaozhong Ji , Chuming Lin , Zhonggan Ding , Ying Tai , Junwei Zhu , Xiaobin Hu , Donghao Luo , Yanhao Ge , Chengjie Wang

Robotic grasping of house-hold objects has made remarkable progress in recent years. Yet, human grasps are still difficult to synthesize realistically. There are several key reasons: (1) the human hand has many degrees of freedom (more than…

Computer Vision and Pattern Recognition · Computer Science 2020-11-30 Korrawe Karunratanakul , Jinlong Yang , Yan Zhang , Michael Black , Krikamol Muandet , Siyu Tang

Socially competent robots should be equipped with the ability to perceive the world that surrounds them and communicate about it in a human-like manner. Representative skills that exhibit such ability include generating image descriptions…

Robotics · Computer Science 2021-02-01 Ting Han , Sina Zarrieß

With the booming of virtual reality (VR) technology, there is a growing need for customized 3D avatars. However, traditional methods for 3D avatar modeling are either time-consuming or fail to retain similarity to the person being modeled.…

Computer Vision and Pattern Recognition · Computer Science 2023-07-06 Chuanyu Pan , Guowei Yang , Taijiang Mu , Yu-Kun Lai

What's the most accurate 3D model of your face you can obtain while sitting at your desk? We attempt to answer this question in our work. High fidelity face reconstructions have so far been limited to either studio settings or through…

Computer Vision and Pattern Recognition · Computer Science 2020-03-20 Shubham Agrawal , Anuj Pahuja , Simon Lucey

Acquiring and annotating sufficient labeled data is crucial in developing accurate and robust learning-based models, but obtaining such data can be challenging in many medical image segmentation tasks. One promising solution is to…

Image and Video Processing · Electrical Eng. & Systems 2023-07-06 Kun Han , Yifeng Xiong , Chenyu You , Pooya Khosravi , Shanlin Sun , Xiangyi Yan , James Duncan , Xiaohui Xie

Natural language plays a critical role in many computer vision applications, such as image captioning, visual question answering, and cross-modal retrieval, to provide fine-grained semantic information. Unfortunately, while human pose is…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Ginger Delmas , Philippe Weinzaepfel , Thomas Lucas , Francesc Moreno-Noguer , Grégory Rogez

This paper aims to achieve the segmentation of any 3D part in a scene based on natural language descriptions, extending beyond traditional object-level 3D scene understanding and addressing both data and methodological challenges. Due to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Hongyu Wu , Pengwan Yang , Yuki M. Asano , Cees G. M. Snoek

Predominant techniques on talking head generation largely depend on 2D information, including facial appearances and motions from input face images. Nevertheless, dense 3D facial geometry, such as pixel-wise depth, plays a critical role in…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Fa-Ting Hong , Li Shen , Dan Xu

Although much progress has been made recently in 3D face reconstruction, most previous work has been devoted to predicting accurate and fine-grained 3D shapes. In contrast, relatively little work has focused on generating high-fidelity face…

Computer Vision and Pattern Recognition · Computer Science 2021-06-16 Xiangnan Yin , Di Huang , Zehua Fu , Yunhong Wang , Liming Chen

We present a minimalistic but effective neural network that computes dense facial correspondences in highly unconstrained RGB images. Our network learns a per-pixel flow and a matchability mask between 2D input photographs of a person and…

Computer Vision and Pattern Recognition · Computer Science 2017-09-05 Ronald Yu , Shunsuke Saito , Haoxiang Li , Duygu Ceylan , Hao Li

Existing text-to-image synthesis methods generally are only applicable to words in the training dataset. However, human faces are so variable to be described with limited words. So this paper proposes the first free-style text-to-face…

Computer Vision and Pattern Recognition · Computer Science 2022-03-30 Jianxin Sun , Qiyao Deng , Qi Li , Muyi Sun , Min Ren , Zhenan Sun

Generating realistic and controllable 3D human avatars is a long-standing challenge, particularly when covering broad attribute ranges such as ethnicity, age, clothing styles, and detailed body shapes. Capturing and annotating large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Yuxuan Xue , Xianghui Xie , Margaret Kostyrko , Gerard Pons-Moll

In recent years, there has been significant progress in 2D generative face models fueled by applications such as animation, synthetic data generation, and digital avatars. However, due to the absence of 3D information, these 2D models often…

Computer Vision and Pattern Recognition · Computer Science 2023-10-30 Aashish Rai , Hiresh Gupta , Ayush Pandey , Francisco Vicente Carrasco , Shingo Jason Takagi , Amaury Aubel , Daeil Kim , Aayush Prakash , Fernando de la Torre

Rapid advancements in text-to-3D generation require robust and scalable evaluation metrics that align closely with human judgment, a need unmet by current metrics such as PSNR and CLIP, which require ground-truth data or focus only on…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Shalini Maiti , Lourdes Agapito , Filippos Kokkinos