中文
相关论文

相关论文: Portrait Video Editing Empowered by Multimodal Gen…

200 篇论文

Video generation models have shown their superior ability to generate photo-realistic video. However, how to accurately control (or edit) the video remains a formidable challenge. The main issues are: 1) how to perform direct and accurate…

图形学 · 计算机科学 2024-07-23 Yufan Deng , Ruida Wang , Yuhao Zhang , Yu-Wing Tai , Chi-Keung Tang

In contrast to the traditional avatar creation pipeline which is a costly process, contemporary generative approaches directly learn the data distribution from photographs. While plenty of works extend unconditional generative models and…

计算机视觉与模式识别 · 计算机科学 2022-10-18 Junshu Tang , Bo Zhang , Binxin Yang , Ting Zhang , Dong Chen , Lizhuang Ma , Fang Wen

One-shot 3D talking portrait generation aims to reconstruct a 3D avatar from an unseen image, and then animate it with a reference video or audio to generate a talking portrait video. The existing methods fail to simultaneously achieve the…

计算机视觉与模式识别 · 计算机科学 2024-03-26 Zhenhui Ye , Tianyun Zhong , Yi Ren , Jiaqi Yang , Weichuang Li , Jiawei Huang , Ziyue Jiang , Jinzheng He , Rongjie Huang , Jinglin Liu , Chen Zhang , Xiang Yin , Zejun Ma , Zhou Zhao

Many recent works have been proposed for face image editing by leveraging the latent space of pretrained GANs. However, few attempts have been made to directly apply them to videos, because 1) they do not guarantee temporal consistency, 2)…

计算机视觉与模式识别 · 计算机科学 2022-06-28 Jiyang Yu , Jingen Liu , Jing Huang , Wei Zhang , Tao Mei

We introduce FactorPortrait, a video diffusion method for controllable portrait animation that enables lifelike synthesis from disentangled control signals of facial expressions, head movement, and camera viewpoints. Given a single portrait…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Jiapeng Tang , Kai Li , Chengxiang Yin , Liuhao Ge , Fei Jiang , Jiu Xu , Matthias Nießner , Christian Häne , Timur Bagautdinov , Egor Zakharov , Peihong Guo

Generating high-quality artistic portrait videos is an important and desirable task in computer graphics and vision. Although a series of successful portrait image toonification models built upon the powerful StyleGAN have been proposed,…

计算机视觉与模式识别 · 计算机科学 2022-10-03 Shuai Yang , Liming Jiang , Ziwei Liu , Chen Change Loy

In the field of portrait video generation, the use of single images to generate portrait videos has become increasingly prevalent. A common approach involves leveraging generative models to enhance adapters for controlled generation.…

计算机视觉与模式识别 · 计算机科学 2024-06-05 Cong Wang , Kuan Tian , Jun Zhang , Yonghang Guan , Feng Luo , Fei Shen , Zhiwei Jiang , Qing Gu , Xiao Han , Wei Yang

We address the challenging problem of generating facial attributes using a single image in an unconstrained pose. In contrast to prior works that largely consider generation on 2D near-frontal images, we propose a GAN-based framework to…

计算机视觉与模式识别 · 计算机科学 2019-07-25 Feng-Ju Chang , Xiang Yu , Ram Nevatia , Manmohan Chandraker

Previous methods have dealt with discrete manipulation of facial attributes such as smile, sad, angry, surprise etc, out of canonical expressions and they are not scalable, operating in single modality. In this paper, we propose a novel…

计算机视觉与模式识别 · 计算机科学 2020-10-07 Jiali Duan , Xiaoyuan Guo , Yuhang Song , Chao Yang , C. -C. Jay Kuo

In this work, we tackle the challenge of enhancing the realism and expressiveness in talking head video generation by focusing on the dynamic and nuanced relationship between audio cues and facial movements. We identify the limitations of…

计算机视觉与模式识别 · 计算机科学 2024-08-09 Linrui Tian , Qi Wang , Bang Zhang , Liefeng Bo

Portrait editing is challenging for existing techniques due to difficulties in preserving subject features like identity. In this paper, we propose a training-based method leveraging auto-generated paired data to learn desired editing while…

计算机视觉与模式识别 · 计算机科学 2024-07-31 Bowei Chen , Tiancheng Zhi , Peihao Zhu , Shen Sang , Jing Liu , Linjie Luo

We observe that recent advances in multimodal foundation models have propelled instruction-driven image generation and editing into a genuinely cross-modal, cooperative regime. Nevertheless, state-of-the-art editing pipelines remain costly:…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Xiaofan Li , Yanpeng Sun , Chenming Wu , Fan Duan , YuAn Wang , Weihao Bo , Yumeng Zhang , Dingkang Liang

Generating 3D models has traditionally been a complex task requiring specialized expertise. While recent advances in generative AI have sought to automate this process, existing methods produce non-editable representation, such as meshes or…

图形学 · 计算机科学 2026-01-21 Fadlullah Raji , Stefano Petrangeli , Matheus Gadelha , Yu Shen , Uttaran Bhattacharya , Gang Wu

Real-time rendering of human head avatars is a cornerstone of many computer graphics applications, such as augmented reality, video games, and films, to name a few. Recent approaches address this challenge with computationally efficient…

计算机视觉与模式识别 · 计算机科学 2024-09-19 Kartik Teotia , Hyeongwoo Kim , Pablo Garrido , Marc Habermann , Mohamed Elgharib , Christian Theobalt

We have recently seen great progress in 3D scene reconstruction through explicit point-based 3D Gaussian Splatting (3DGS), notable for its high quality and fast rendering speed. However, reconstructing dynamic scenes such as complex human…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Chao Zhang , Yifeng Zhou , Shuheng Wang , Wenfa Li , Degang Wang , Yi Xu , Shaohui Jiao

Face reenactment methods attempt to restore and re-animate portrait videos as realistically as possible. Existing methods face a dilemma in quality versus controllability: 2D GAN-based methods achieve higher image quality but suffer in…

计算机视觉与模式识别 · 计算机科学 2023-05-02 Lizhen Wang , Xiaochen Zhao , Jingxiang Sun , Yuxiang Zhang , Hongwen Zhang , Tao Yu , Yebin Liu

Image-to-Video generation (I2V) animates a static image into a temporally coherent video sequence following textual instructions, yet preserving fine-grained object identity under changing viewpoints remains a persistent challenge. Unlike…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Mingyang Wu , Ashirbad Mishra , Soumik Dey , Shuo Xing , Naveen Ravipati , Hansi Wu , Binbin Li , Zhengzhong Tu

Recently, Generative Adversarial Networks (GANs)} have been widely used for portrait image generation. However, in the latent space learned by GANs, different attributes, such as pose, shape, and texture style, are generally entangled,…

计算机视觉与模式识别 · 计算机科学 2021-08-09 Anpei Chen , Ruiyang Liu , Ling Xie , Zhang Chen , Hao Su , Jingyi Yu

Existing methods like Neural Radiation Fields (NeRF) and 3D Gaussian Splatting (3DGS) have made significant strides in facial attribute control such as facial animation and components editing, yet they struggle with fine-grained…

计算机视觉与模式识别 · 计算机科学 2024-08-21 Pinxin Liu , Luchuan Song , Daoan Zhang , Hang Hua , Yunlong Tang , Huaijin Tu , Jiebo Luo , Chenliang Xu

Synthesizing consistent and photorealistic 3D scenes is an open problem in computer vision. Video diffusion models generate impressive videos but cannot directly synthesize 3D representations, i.e., lack 3D consistency in the generated…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Katja Schwarz , Norman Mueller , Peter Kontschieder