中文
相关论文

相关论文: Loki: Representation over Architecture for Diffusi…

200 篇论文

Audio-driven portrait animation aims to synthesize realistic and natural talking head videos from an input audio signal and a single reference image. While existing methods achieve high-quality results by leveraging high-dimensional…

图形学 · 计算机科学 2026-02-27 Fangyu Du , Taiqing Li , Qian Qiao , Tan Yu , Ziwei Zhang , Dingcheng Zhen , Xu Jia , Yang Yang , Shunshun Yin , Siyuan Liu

Generating highly dynamic and photorealistic portrait animations driven by audio and skeletal motion remains challenging due to the need for precise lip synchronization, natural facial expressions, and high-fidelity body motion dynamics. We…

计算机视觉与模式识别 · 计算机科学 2025-12-02 Jiahao Cui , Yan Chen , Mingwang Xu , Hanlin Shang , Yuxuan Chen , Yun Zhan , Zilong Dong , Yao Yao , Jingdong Wang , Siyu Zhu

Pose-driven full-body avatars built on neural rendering produce high-quality novel views of a captured subject. Yet loose clothing and other dynamic elements deform in ways pose alone cannot explain: the same pose can correspond to many…

计算机视觉与模式识别 · 计算机科学 2026-05-21 Shichong Peng , Chengxiang Yin , Fei Jiang , Zhongshi Jiang , Lingchen Yang , Qingyang Tan , Amin Jourabloo , Jason Saragih , Ke Li , Christian Häne

Animating stylized avatars with dynamic poses and expressions has attracted increasing attention for its broad range of applications. Previous research has made significant progress by training controllable generative models to synthesize…

Blind face restoration is a highly ill-posed problem due to the lack of necessary context. Although existing methods produce high-quality outputs, they often fail to faithfully preserve the individual's identity. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Siyu Liu , Zheng-Peng Duan , Jia OuYang , Jiayi Fu , Hyunhee Park , Zikun Liu , Chun-Le Guo , Chongyi Li

Imitation learning provides an efficient way to teach robots dexterous skills; however, learning complex skills robustly and generalizablely usually consumes large amounts of human demonstrations. To tackle this challenging problem, we…

机器人学 · 计算机科学 2024-09-30 Yanjie Ze , Gu Zhang , Kangning Zhang , Chenyuan Hu , Muhan Wang , Huazhe Xu

Facial appearance editing is crucial for digital avatars, AR/VR, and personalized content creation, driving realistic user experiences. However, preserving identity with generative models is challenging, especially in scenarios with limited…

计算机视觉与模式识别 · 计算机科学 2025-03-10 MD Wahiduzzaman Khan , Mingshan Jia , Xiaolin Zhang , En Yu , Caifeng Shan , Kaska Musial-Gabrys

In human-centric content generation, the pre-trained text-to-image models struggle to produce user-wanted portrait images, which retain the identity of individuals while exhibiting diverse expressions. This paper introduces our efforts…

计算机视觉与模式识别 · 计算机科学 2024-04-09 Renshuai Liu , Bowen Ma , Wei Zhang , Zhipeng Hu , Changjie Fan , Tangjie Lv , Yu Ding , Xuan Cheng

Diffusion-based methods have demonstrated remarkable capabilities in generating a diverse array of high-quality images, sparking interests for styled avatars, virtual try-on, and more. Previous methods use the same reference image as the…

计算机视觉与模式识别 · 计算机科学 2024-10-24 Haoran Tang , Jieren Deng , Zhihong Pan , Hao Tian , Pratik Chaudhari , Xin Zhou

Millions of images of human faces are captured every single day; but these photographs portray the likeness of an individual with a fixed pose, expression, and appearance. Portrait image animation enables the post-capture adjustment of…

计算机视觉与模式识别 · 计算机科学 2022-03-28 Connor Z. Lin , David B. Lindell , Eric R. Chan , Gordon Wetzstein

We propose a method to transfer pose and expression between face images. Given a source and target face portrait, the model produces an output image in which the pose and expression of the source face image are transferred onto the target…

计算机视觉与模式识别 · 计算机科学 2025-04-18 Petr Jahoda , Jan Cech

We present techniques for improving performance driven facial animation, emotion recognition, and facial key-point or landmark prediction using learned identity invariant representations. Established approaches to these problems can work…

计算机视觉与模式识别 · 计算机科学 2016-05-24 David Rim , Sina Honari , Md Kamrul Hasan , Chris Pal

Large-scale pre-trained diffusion models have exhibited remarkable capabilities in diverse video generations. Given a set of video clips of the same motion concept, the task of Motion Customization is to adapt existing text-to-video…

计算机视觉与模式识别 · 计算机科学 2023-10-13 Rui Zhao , Yuchao Gu , Jay Zhangjie Wu , David Junhao Zhang , Jiawei Liu , Weijia Wu , Jussi Keppo , Mike Zheng Shou

The development of diffusion-based generative models over the past decade has largely proceeded independently of progress in representation learning. These diffusion models typically rely on regression-based objectives and generally lack…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Runqian Wang , Kaiming He

Diffusion models (DMs) have achieved state-of-the-art results for image synthesis tasks as well as density estimation. Applied in the latent space of a powerful pretrained autoencoder (LDM), their immense computational requirements can be…

计算机视觉与模式识别 · 计算机科学 2022-10-21 Jeremias Traub

Text-to-image models (T2I) such as StableDiffusion have been used to generate high quality images of people. However, due to the random nature of the generation process, the person has a different appearance e.g. pose, face, and clothing,…

计算机视觉与模式识别 · 计算机科学 2023-12-19 Soon Yau Cheong , Armin Mustafa , Andrew Gilbert

Face reenactment aims to generate realistic talking head videos by transferring motion from a driving video to a static source image while preserving the source identity. Although existing methods based on either implicit or explicit…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Mingtao Guo , Guanyu Xing , Yanci Zhang , Yanli Liu

Pose-guided person image generation usually involves using paired source-target images to supervise the training, which significantly increases the data preparation effort and limits the application of the models. To deal with this problem,…

计算机视觉与模式识别 · 计算机科学 2021-04-12 Tianxiang Ma , Bo Peng , Wei Wang , Jing Dong

We propose a method to learn, even using a dataset where objects appear only in sparsely sampled views (e.g. Pix3D), the ability to synthesize a pose trajectory for an arbitrary reference image. This is achieved with a cross-modal pose…

计算机视觉与模式识别 · 计算机科学 2021-05-04 Bo Liu , Mandar Dixit , Roland Kwitt , Gang Hua , Nuno Vasconcelos

The rapid progress of Deepfake technology has made face swapping highly realistic, raising concerns about the malicious use of fabricated facial content. Existing methods often struggle to generalize to unseen domains due to the diverse…

计算机视觉与模式识别 · 计算机科学 2024-10-08 Ke Sun , Shen Chen , Taiping Yao , Hong Liu , Xiaoshuai Sun , Shouhong Ding , Rongrong Ji