English
Related papers

Related papers: Disentangling 3D from Large Vision-Language Models…

200 papers

Diffusion-based generative models have exhibited powerful generative performance in recent years. However, as many attributes exist in the data distribution and owing to several limitations of sharing the model parameters across all levels…

Audio and Speech Processing · Electrical Eng. & Systems 2023-05-26 Ha-Yeong Choi , Sang-Hoon Lee , Seong-Whan Lee

Recent advancements in implicit 3D representations and generative models have markedly propelled the field of 3D object generation forward. However, it remains a significant challenge to accurately model geometries with defined sharp…

Graphics · Computer Science 2024-01-17 Zeqing Yuan , Haoxuan Lan , Qiang Zou , Junbo Zhao

Personalized image generation has emerged as a promising direction in multimodal content creation. It aims to synthesize images tailored to individual style preferences (e.g., color schemes, character appearances, layout) and semantic…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Yiyan Xu , Wuqiang Zheng , Wenjie Wang , Fengbin Zhu , Xinting Hu , Yang Zhang , Fuli Feng , Tat-Seng Chua

Confounding bias is a crucial problem when applying machine learning to practice, especially in clinical practice. We consider the problem of learning representations independent to multiple biases. In literature, this is mostly solved by…

Computer Vision and Pattern Recognition · Computer Science 2021-06-30 Xianjing Liu , Bo Li , Esther Bron , Wiro Niessen , Eppo Wolvius , Gennady Roshchupkin

We introduce FaceGPT, a self-supervised learning framework for Large Vision-Language Models (VLMs) to reason about 3D human faces from images and text. Typical 3D face reconstruction methods are specialized algorithms that lack semantic…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Haoran Wang , Mohit Mendiratta , Christian Theobalt , Adam Kortylewski

A crucial problem in learning disentangled image representations is controlling the degree of disentanglement during image editing, while preserving the identity of objects. In this work, we propose a simple yet effective model with the…

Machine Learning · Computer Science 2019-12-30 Zengjie Song , Oluwasanmi Koyejo , Jiangshe Zhang

We propose a framework, called LiftedGAN, that disentangles and lifts a pre-trained StyleGAN2 for 3D-aware face generation. Our model is "3D-aware" in the sense that it is able to (1) disentangle the latent space of StyleGAN2 into texture,…

Computer Vision and Pattern Recognition · Computer Science 2021-04-20 Yichun Shi , Divyansh Aggarwal , Anil K. Jain

Learning visual representations with interpretable features, i.e., disentangled representations, remains a challenging problem. Existing methods demonstrate some success but are hard to apply to large-scale vision datasets like ImageNet. In…

Machine Learning · Computer Science 2023-06-01 Lilian Ngweta , Subha Maity , Alex Gittens , Yuekai Sun , Mikhail Yurochkin

Learning disentangled representations of data is a fundamental problem in artificial intelligence. Specifically, disentangled latent representations allow generative models to control and compose the disentangled factors in the synthesis…

Computer Vision and Pattern Recognition · Computer Science 2020-10-20 Yotam Nitzan , Amit Bermano , Yangyan Li , Daniel Cohen-Or

Despite recent advancements in 3D generation methods, achieving controllability still remains a challenging issue. Current approaches utilizing score-distillation sampling are hindered by laborious procedures that consume a significant…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Hongbin Xu , Weitao Chen , Zhipeng Zhou , Feng Xiao , Baigui Sun , Mike Zheng Shou , Wenxiong Kang

In this work we introduce Lifting Autoencoders, a generative 3D surface-based model of object categories. We bring together ideas from non-rigid structure from motion, image formation, and morphable models to learn a controllable, geometric…

Computer Vision and Pattern Recognition · Computer Science 2019-04-29 Mihir Sahasrabudhe , Zhixin Shu , Edward Bartrum , Riza Alp Guler , Dimitris Samaras , Iasonas Kokkinos

Designing realistic digital humans is extremely complex. Most data-driven generative models used to simplify the creation of their underlying geometric shape do not offer control over the generation of local shape attributes. In this paper,…

Computer Vision and Pattern Recognition · Computer Science 2023-04-05 Simone Foti , Bongjin Koo , Danail Stoyanov , Matthew J. Clarkson

We explore different design choices for injecting noise into generative adversarial networks (GANs) with the goal of disentangling the latent space. Instead of traditional approaches, we propose feeding multiple noise codes through separate…

Computer Vision and Pattern Recognition · Computer Science 2020-05-06 Yazeed Alharbi , Peter Wonka

Recent strides in Text-to-3D techniques have been propelled by distilling knowledge from powerful large text-to-image diffusion models (LDMs). Nonetheless, existing Text-to-3D approaches often grapple with challenges such as…

Computer Vision and Pattern Recognition · Computer Science 2023-08-23 Yiwen Chen , Chi Zhang , Xiaofeng Yang , Zhongang Cai , Gang Yu , Lei Yang , Guosheng Lin

Large Vision-Language Models (LVLMs) have demonstrated remarkable capabilities in processing both visual and textual information. However, the critical challenge of alignment between visual and textual representations is not fully…

Computer Vision and Pattern Recognition · Computer Science 2025-09-24 Dong Shu , Haiyan Zhao , Jingyu Hu , Weiru Liu , Ali Payani , Lu Cheng , Mengnan Du

Learning a disentangled, interpretable, and structured latent representation in 3D generative models of faces and bodies is still an open problem. The problem is particularly acute when control over identity features is required. In this…

Computer Vision and Pattern Recognition · Computer Science 2022-03-25 Simone Foti , Bongjin Koo , Danail Stoyanov , Matthew J. Clarkson

Near-range portrait photographs often contain perspective distortion artifacts that bias human perception and challenge both facial recognition and reconstruction techniques. We present the first deep learning based approach to remove such…

Computer Vision and Pattern Recognition · Computer Science 2019-05-21 Yajie Zhao , Zeng Huang , Tianye Li , Weikai Chen , Chloe LeGendre , Xinglei Ren , Jun Xing , Ari Shapiro , Hao Li

Disentangled visual representations have largely been studied with generative models such as Variational AutoEncoders (VAEs). While prior work has focused on generative methods for disentangled representation learning, these approaches do…

Computer Vision and Pattern Recognition · Computer Science 2021-08-17 Andrea Burns , Aaron Sarna , Dilip Krishnan , Aaron Maschinot

Recent advances in generative adversarial networks (GANs) have led to remarkable achievements in face image synthesis. While methods that use style-based GANs can generate strikingly photorealistic face images, it is often difficult to…

Computer Vision and Pattern Recognition · Computer Science 2021-11-02 Safa C. Medin , Bernhard Egger , Anoop Cherian , Ye Wang , Joshua B. Tenenbaum , Xiaoming Liu , Tim K. Marks

Face image manipulation via three-dimensional guidance has been widely applied in various interactive scenarios due to its semantically-meaningful understanding and user-friendly controllability. However, existing 3D-morphable-model-based…

Computer Vision and Pattern Recognition · Computer Science 2022-03-01 Can Wang , Menglei Chai , Mingming He , Dongdong Chen , Jing Liao