中文
相关论文

相关论文: Towards High-Fidelity Text-Guided 3D Face Generati…

200 篇论文

We propose ID-to-3D, a method to generate identity- and text-guided 3D human heads with disentangled expressions, starting from even a single casually captured in-the-wild image of a subject. The foundation of our approach is anchored in…

计算机视觉与模式识别 · 计算机科学 2024-05-29 Francesca Babiloni , Alexandros Lattas , Jiankang Deng , Stefanos Zafeiriou

Facial images have extensive practical applications. Although the current large-scale text-image diffusion models exhibit strong generation capabilities, it is challenging to generate the desired facial images using only text prompt. Image…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Dawei Dai , Mingming Jia , Yinxiu Zhou , Hang Xing , Chenghang Li

The generation of high-quality, animatable 3D head avatars from text has enormous potential in content creation applications such as games, movies, and embodied virtual assistants. Current text-to-3D generation methods typically combine…

计算机视觉与模式识别 · 计算机科学 2025-04-23 Yiqian Wu , Malte Prinzler , Xiaogang Jin , Siyu Tang

We present a new multi-modal face image generation method that converts a text prompt and a visual input, such as a semantic mask or scribble map, into a photo-realistic face image. To do this, we combine the strengths of Generative…

计算机视觉与模式识别 · 计算机科学 2024-05-08 Jihyun Kim , Changjae Oh , Hoseok Do , Soohyun Kim , Kwanghoon Sohn

Capitalizing on the recent advances in image generation models, existing controllable face image synthesis methods are able to generate high-fidelity images with some levels of controllability, e.g., controlling the shapes, expressions,…

计算机视觉与模式识别 · 计算机科学 2023-10-31 Keqiang Sun , Shangzhe Wu , Ning Zhang , Zhaoyang Huang , Quan Wang , Hongsheng Li

We train a feed-forward text-to-3D diffusion generator for human characters using only single-view 2D data for supervision. Existing 3D generative models cannot yet match the fidelity of image or video generative models. State-of-the-art 3D…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Souhaib Attaiki , Paul Guerrero , Duygu Ceylan , Niloy J. Mitra , Maks Ovsjanikov

While recent generative models for 2D images achieve impressive visual results, they clearly lack the ability to perform 3D reasoning. This heavily restricts the degree of control over generated objects as well as the possible applications…

计算机视觉与模式识别 · 计算机科学 2020-10-26 Dario Pavllo , Graham Spinks , Thomas Hofmann , Marie-Francine Moens , Aurelien Lucchi

Acquiring detailed 3D scenes typically demands costly equipment, multi-view data, or labor-intensive modeling. Therefore, a lightweight alternative, generating complex 3D scenes from a single top-down image, plays an essential role in…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Kaizhi Zheng , Ruijian Zha , Zishuo Xu , Jing Gu , Jie Yang , Xin Eric Wang

Face editing modifies the appearance of face, which plays a key role in customization and enhancement of personal images. Although much work have achieved remarkable success in text-driven face editing, they still face significant…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Xin Zhang , Siting Huang , Xiangyang Luo , Yifan Xie , Weijiang Yu , Heng Chang , Fei Ma , Fei Yu

Generating 3D models has traditionally been a complex task requiring specialized expertise. While recent advances in generative AI have sought to automate this process, existing methods produce non-editable representation, such as meshes or…

图形学 · 计算机科学 2026-01-21 Fadlullah Raji , Stefano Petrangeli , Matheus Gadelha , Yu Shen , Uttaran Bhattacharya , Gang Wu

Talking head generation is to generate video based on a given source identity and target motion. However, current methods face several challenges that limit the quality and controllability of the generated videos. First, the generated face…

计算机视觉与模式识别 · 计算机科学 2023-11-03 Yue Gao , Yuan Zhou , Jinglu Wang , Xiao Li , Xiang Ming , Yan Lu

We develop an approach for text-to-image generation that embraces additional retrieval images, driven by a combination of implicit visual guidance loss and generative objectives. Unlike most existing text-to-image generation methods which…

计算机视觉与模式识别 · 计算机科学 2022-08-19 Xin Yuan , Zhe Lin , Jason Kuen , Jianming Zhang , John Collomosse

A 3D avatar typically has one of six cardinal facial expressions. To simulate realistic emotional variation, we should be able to render a facial transition between two arbitrary expressions. This study presents a new framework for…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Anh H. Vo , Tae-Seok Kim , Hulin Jin , Soo-Mi Choi , Yong-Guk Kim

Generative Adversarial Networks (GANs) have revolutionized image synthesis through many applications like face generation, photograph editing, and image super-resolution. Image synthesis using GANs has predominantly been uni-modal, with few…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Rohan Wadhawan , Tanuj Drall , Shubham Singh , Shampa Chakraverty

We propose a novel technique for adding geometric details to an input coarse 3D mesh guided by a text prompt. Our method is composed of three stages. First, we generate a single-view RGB image conditioned on the input coarse geometry and…

计算机视觉与模式识别 · 计算机科学 2024-09-12 Yun-Chun Chen , Selena Ling , Zhiqin Chen , Vladimir G. Kim , Matheus Gadelha , Alec Jacobson

In this work, we propose TediGAN, a novel framework for multi-modal image generation and manipulation with textual descriptions. The proposed method consists of three components: StyleGAN inversion module, visual-linguistic similarity…

计算机视觉与模式识别 · 计算机科学 2021-03-30 Weihao Xia , Yujiu Yang , Jing-Hao Xue , Baoyuan Wu

Vivid talking face generation holds immense potential applications across diverse multimedia domains, such as film and game production. While existing methods accurately synchronize lip movements with input audio, they typically ignore…

计算机视觉与模式识别 · 计算机科学 2024-06-13 Jiadong Liang , Feng Lu

Text-to-3D generation has made remarkable progress recently, particularly with methods based on Score Distillation Sampling (SDS) that leverages pre-trained 2D diffusion models. While the usage of classifier-free guidance is well…

计算机视觉与模式识别 · 计算机科学 2023-11-01 Xin Yu , Yuan-Chen Guo , Yangguang Li , Ding Liang , Song-Hai Zhang , Xiaojuan Qi

Text-guided human body animation has advanced rapidly, yet facial animation lags due to the scarcity of well-annotated, text-paired facial corpora. To close this gap, we leverage foundation generative models to synthesize a large, balanced…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Luchuan Song , Pinxin Liu , Haiyang Liu , Zhenchao Jin , Yolo Yunlong Tang , Zichong Xu , Susan Liang , Jing Bi , Jason J Corso , Chenliang Xu

This paper presents a new text-guided technique for generating 3D shapes. The technique leverages a hybrid 3D shape representation, namely EXIM, combining the strengths of explicit and implicit representations. Specifically, the explicit…

计算机视觉与模式识别 · 计算机科学 2023-12-01 Zhengzhe Liu , Jingyu Hu , Ka-Hei Hui , Xiaojuan Qi , Daniel Cohen-Or , Chi-Wing Fu