中文
相关论文

相关论文: Semantic Latent Decomposition with Normalizing Flo…

200 篇论文

Existing methods for face image manipulation generally focus on editing the expression, changing some predefined attributes, or applying different filters. However, users lack the flexibility of controlling the shapes of different semantic…

计算机视觉与模式识别 · 计算机科学 2019-05-07 Sen-Zhe Xu , Hao-Zhi Huang , Shi-Min Hu , Wei Liu

Speech is a scalable and non-invasive biomarker for early mental health screening. However, widely used depression datasets like DAIC-WOZ exhibit strong coupling between linguistic sentiment and diagnostic labels, encouraging models to…

计算与语言 · 计算机科学 2026-01-05 Yuxin Li , Xiangyu Zhang , Yifei Li , Zhiwei Guo , Haoyang Zhang , Eng Siong Chng , Cuntai Guan

Semantic watermarking techniques for latent diffusion models (LDMs) are robust against regeneration attacks, but often suffer from detection performance degradation due to the loss of frequency integrity. To tackle this problem, we propose…

计算机视觉与模式识别 · 计算机科学 2025-09-10 Sung Ju Lee , Nam Ik Cho

In this pioneering study, we introduce StyleWallfacer, a groundbreaking unified training and inference framework, which not only addresses various issues encountered in the style transfer process of traditional methods but also unifies the…

计算机视觉与模式识别 · 计算机科学 2025-06-19 Gary Song Yan , Yusen Zhang , Jinyu Zhao , Hao Zhang , Zhangping Yang , Guanye Xiong , Yanfei Liu , Tao Zhang , Yujie He , Siyuan Tian , Yao Gou , Min Li

Unconstrained Image generation with high realism is now possible using recent Generative Adversarial Networks (GANs). However, it is quite challenging to generate images with a given set of attributes. Recent methods use style-based GAN…

计算机视觉与模式识别 · 计算机科学 2022-07-21 Rishubh Parihar , Ankit Dhiman , Tejan Karmali , R. Venkatesh Babu

We propose a local adversarial disentangling network (LADN) for facial makeup and de-makeup. Central to our method are multiple and overlapping local adversarial discriminators in a content-style disentangling network for achieving local…

计算机视觉与模式识别 · 计算机科学 2019-08-12 Qiao Gu , Guanzhi Wang , Mang Tik Chiu , Yu-Wing Tai , Chi-Keung Tang

This paper presents a new approach for the detection of fake videos, based on the analysis of style latent vectors and their abnormal behavior in temporal changes in the generated videos. We discovered that the generated facial videos…

计算机视觉与模式识别 · 计算机科学 2024-05-21 Jongwook Choi , Taehoon Kim , Yonghyun Jeong , Seungryul Baek , Jongwon Choi

Flow matching and diffusion models have shown impressive results in text-to-image generation, producing photorealistic images through an iterative denoising process. A common strategy to speed up synthesis is to perform early denoising at…

计算机视觉与模式识别 · 计算机科学 2025-07-23 Jyun-Ze Tang , Chih-Fan Hsu , Jeng-Lin Li , Ming-Ching Chang , Wei-Chao Chen

Interactive facial image manipulation attempts to edit single and multiple face attributes using a photo-realistic face and/or semantic mask as input. In the absence of the photo-realistic image (only sketch/mask available), previous…

计算机视觉与模式识别 · 计算机科学 2021-10-12 Yan Yang , Md Zakir Hossain , Tom Gedeon , Shafin Rahman

Realistic generative face video synthesis has long been a pursuit in both computer vision and graphics community. However, existing face video generation methods tend to produce low-quality frames with drifted facial identities and…

计算机视觉与模式识别 · 计算机科学 2022-08-17 Haonan Qiu , Yuming Jiang , Hang Zhou , Wayne Wu , Ziwei Liu

Drawing upon StyleGAN's expressivity and disentangled latent space, existing 2D approaches employ textual prompting to edit facial images with different attributes. In contrast, 3D-aware approaches that generate faces at different target…

计算机视觉与模式识别 · 计算机科学 2024-07-25 Amandeep Kumar , Muhammad Awais , Sanath Narayan , Hisham Cholakkal , Salman Khan , Rao Muhammad Anwer

Morphing is a long-standing problem in vision and computer graphics, requiring a time-dependent warping for feature alignment and a blending for smooth interpolation. Recently, multilayer perceptrons (MLPs) have been explored as implicit…

Normalizing flows (NFs) provide exact likelihoods and deterministic invertible sampling, but have historically lagged behind diffusion models for large-scale image generation. We identify a key obstacle: NFs are required to learn a single…

计算机视觉与模式识别 · 计算机科学 2026-05-26 Longtao Jiang , Jianmin Bao , Zhendong Wang , Xin Tao , Pengfei Wan , Zhihui Li , Xiaojun Chang

Standard flow matching scales well but typically relies on an unstructured source distribution, limiting its ability to learn interpretable latent structure. Latent-variable models, by contrast, capture structure but often sacrifice…

机器学习 · 计算机科学 2026-05-11 Xavier Sumba , Carles Balsells-Rodas , Yingzhen Li

Although many deep-learning-based super-resolution approaches have been proposed in recent years, because no ground truth is available in the inference stage, few can quantify the errors and uncertainties of the super-resolved results. For…

图像与视频处理 · 电气工程与系统科学 2023-08-10 Jingyi Shen , Han-Wei Shen

Recently, there has been a surge of diverse methods for performing image editing by employing pre-trained unconditional generators. Applying these methods on real images, however, remains a challenge, as it necessarily requires the…

计算机视觉与模式识别 · 计算机科学 2021-02-05 Omer Tov , Yuval Alaluf , Yotam Nitzan , Or Patashnik , Daniel Cohen-Or

Existing methods for the 4D reconstruction of general, non-rigidly deforming objects focus on novel-view synthesis and neglect correspondences. However, time consistency enables advanced downstream tasks like 3D editing, motion analysis, or…

计算机视觉与模式识别 · 计算机科学 2023-08-17 Edith Tretschk , Vladislav Golyanik , Michael Zollhoefer , Aljaz Bozic , Christoph Lassner , Christian Theobalt

Recent advances in inverse problem solving have increasingly adopted flow priors over diffusion models due to their ability to construct straight probability paths from noise to data, thereby enhancing efficiency in both training and…

计算机视觉与模式识别 · 计算机科学 2025-11-11 Hossein Askari , Yadan Luo , Hongfu Sun , Fred Roosta

We propose a novel, vision-only object-level SLAM framework for automotive applications representing 3D shapes by implicit signed distance functions. Our key innovation consists of augmenting the standard neural representation by a…

计算机视觉与模式识别 · 计算机科学 2025-03-17 Li Cui , Yang Ding , Richard Hartley , Zirui Xie , Laurent Kneip , Zhenghua Yu

Flow matching models have recently emerged as an efficient alternative to diffusion, especially for text-guided image generation and editing, offering faster inference through continuous-time dynamics. However, existing flow-based editors…

计算机视觉与模式识别 · 计算机科学 2026-02-11 Carmine Zaccagnino , Fabio Quattrini , Enis Simsar , Marta Tintoré Gazulla , Rita Cucchiara , Alessio Tonioni , Silvia Cascianelli