English
Related papers

Related papers: DVI: Disentangling Semantic and Visual Identity fo…

200 papers

Learning disentangled representations in sequential data is a key goal in deep learning, with broad applications in vision, audio, and time series. While real-world data involves multiple interacting semantic factors over time, prior work…

Machine Learning · Computer Science 2025-10-28 Tal Barami , Nimrod Berman , Ilan Naiman , Amos H. Hason , Rotem Ezra , Omri Azencot

We propose a novel approach for disentangling visual and semantic features from the backbones of pre-trained diffusion models, enabling visual correspondence in a manner analogous to the well-established semantic correspondence. While…

Computer Vision and Pattern Recognition · Computer Science 2025-09-29 Abdelrahman Eldesokey , Aleksandar Cvejic , Bernard Ghanem , Peter Wonka

Building facial analysis systems that generalize to extreme variations in lighting and facial expressions is a challenging problem that can potentially be alleviated using natural-looking synthetic data. Towards that, we propose LEGAN, a…

Computer Vision and Pattern Recognition · Computer Science 2021-06-22 Sandipan Banerjee , Ajjen Joshi , Prashant Mahajan , Sneha Bhattacharya , Survi Kyal , Taniya Mishra

We introduce the first zero-shot approach for Video Semantic Segmentation (VSS) based on pre-trained diffusion models. A growing research direction attempts to employ diffusion models to perform downstream vision tasks by exploiting their…

Computer Vision and Pattern Recognition · Computer Science 2024-05-28 Qian Wang , Abdelrahman Eldesokey , Mohit Mendiratta , Fangneng Zhan , Adam Kortylewski , Christian Theobalt , Peter Wonka

Recent progress in diffusion-based visual generation has largely relied on latent diffusion models with variational autoencoders (VAEs). While effective for high-fidelity synthesis, this VAE+diffusion paradigm suffers from limited training…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Minglei Shi , Haolin Wang , Wenzhao Zheng , Ziyang Yuan , Xiaoshi Wu , Xintao Wang , Pengfei Wan , Jie Zhou , Jiwen Lu

We propose a method to disentangle linear-encoded facial semantics from StyleGAN without external supervision. The method derives from linear regression and sparse representation learning concepts to make the disentangled latent…

Computer Vision and Pattern Recognition · Computer Science 2021-04-01 Yutong Zheng , Yu-Kai Huang , Ran Tao , Zhiqiang Shen , Marios Savvides

Diffusion Transformers (DiTs) have recently achieved remarkable success in text-guided image generation. In image editing, DiTs project text and image inputs to a joint latent space, from which they decode and synthesize new images.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Zitao Shuai , Chenwei Wu , Zhengxu Tang , Bowen Song , Liyue Shen

Machine-learning models applied to skin images often have degraded performance when the skin colour captured in images (SCCI) differs between training and deployment. These discrepancies arise from a combination of entangled environmental…

Image and Video Processing · Electrical Eng. & Systems 2026-02-27 Wenbo Yang , Eman Rezk , Walaa M. Moursi , Zhou Wang

Recent studies on facial expression editing have obtained very promising progress. On the other hand, existing methods face the constraint of requiring a large amount of expression labels which are often expensive and time-consuming to…

Computer Vision and Pattern Recognition · Computer Science 2020-07-20 Rongliang Wu , Shijian Lu

In the last decade, numerous supervised deep learning approaches requiring large amounts of labeled data have been proposed for visual-inertial odometry (VIO) and depth map estimation. To overcome the data limitation, self-supervised…

Computer Vision and Pattern Recognition · Computer Science 2020-07-24 Yasin Almalioglu , Mehmet Turan , Alp Eren Sari , Muhamad Risqi U. Saputra , Pedro P. B. de Gusmão , Andrew Markham , Niki Trigoni

Diffusion models have established the state-of-the-art in text-to-image generation, but their performance often relies on a diffusion prior network to translate text embeddings into the visual manifold for easier decoding. These priors are…

Computer Vision and Pattern Recognition · Computer Science 2025-12-30 Samuele Dell'Erba , Andrew D. Bagdanov

Text-to-image diffusion models have demonstrated remarkable capabilities in generating high-quality images, yet their tendency to reproduce undesirable concepts, such as NSFW content, copyrighted styles, or specific objects, poses growing…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Zhiqi Zhang , Xinhao Zhong , Yi Sun , Shuoyang Sun , Bin Chen , Shu-Tao Xia , Xuan Wang

Face anonymization aims to conceal the visual identity of a face to safeguard the individual's privacy. Traditional methods like blurring and pixelation can largely remove identifying features, but these techniques significantly degrade…

Computer Vision and Pattern Recognition · Computer Science 2025-01-17 Lin Yuan , Kai Liang , Xiong Li , Tao Wu , Nannan Wang , Xinbo Gao

Contrastive pre-trained vision-language models, such as CLIP, demonstrate strong generalization abilities in zero-shot classification by leveraging embeddings extracted from image and text encoders. This paper aims to robustly fine-tune…

Computer Vision and Pattern Recognition · Computer Science 2025-11-14 Satoshi Suzuki , Shin'ya Yamaguchi , Shoichiro Takeda , Taiga Yamane , Naoki Makishima , Naotaka Kawata , Mana Ihori , Tomohiro Tanaka , Shota Orihashi , Ryo Masumura

A variational autoencoder (VAE) is a probabilistic machine learning framework for posterior inference that projects an input set of high-dimensional data to a lower-dimensional, latent space. The latent space learned with a VAE offers…

Machine Learning · Computer Science 2022-11-16 Rafael Pastrana

Concept personalization methods enable large text-to-image models to learn specific subjects (e.g., objects/poses/3D models) and synthesize renditions in new contexts. Given that the image references are highly biased towards visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 You Wu , Kean Liu , Xiaoyue Mi , Fan Tang , Juan Cao , Jintao Li

Traditional studies on voice conversion (VC) have made progress with parallel training data and known speakers. Good voice conversion quality is obtained by exploring better alignment modules or expressive mapping functions. In this study,…

Audio and Speech Processing · Electrical Eng. & Systems 2022-04-01 Jiachen Lian , Chunlei Zhang , Dong Yu

The use of well-disentangled representations offers many advantages for downstream tasks, e.g. an increased sample efficiency, or better interpretability. However, the quality of disentangled interpretations is often highly dependent on the…

Machine Learning · Computer Science 2023-03-03 Benjamin Estermann , Roger Wattenhofer

The ability to recognize objects despite there being differences in appearance, known as Core Object Recognition, forms a critical part of human perception. While it is understood that the brain accomplishes Core Object Recognition through…

Machine Learning · Computer Science 2020-05-15 Harshvardhan Sikka

Privacy of machine learning models is one of the remaining challenges that hinder the broad adoption of Artificial Intelligent (AI). This paper considers this problem in the context of image datasets containing faces. Anonymization of such…

Computer Vision and Pattern Recognition · Computer Science 2022-12-29 Minh-Ha Le , Niklas Carlsson