English
Related papers

Related papers: Improving Style-Content Disentanglement in Image-t…

200 papers

We contribute an unsupervised method that effectively learns disentangled content and style representations from sequences of observations. Unlike most disentanglement algorithms that rely on domain-specific labels or knowledge, our method…

Machine Learning · Computer Science 2025-03-18 Yuxuan Wu , Ziyu Wang , Bhiksha Raj , Gus Xia

Text-to-image diffusion models excel at generating high-quality images from natural language descriptions but often fail to preserve subject consistency across multiple outputs, limiting their use in visual storytelling. Existing approaches…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Shangxun Li , Youngjung Uh

The diffusion-based text-to-image model harbors immense potential in transferring reference style. However, current encoder-based approaches significantly impair the text controllability of text-to-image models while transferring styles. In…

Computer Vision and Pattern Recognition · Computer Science 2024-03-13 Tianhao Qi , Shancheng Fang , Yanze Wu , Hongtao Xie , Jiawei Liu , Lang Chen , Qian He , Yongdong Zhang

We propose a novel spatially-correlative loss that is simple, efficient and yet effective for preserving scene structure consistency while supporting large appearance changes during unpaired image-to-image (I2I) translation. Previous…

Computer Vision and Pattern Recognition · Computer Science 2021-04-05 Chuanxia Zheng , Tat-Jen Cham , Jianfei Cai

Text-to-image diffusion models have emerged as powerful tools for high-quality image generation and editing. Many existing approaches rely on text prompts as editing guidance. However, these methods are constrained by the need for manual…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Yuanyuan Chang , Yinghua Yao , Tao Qin , Mengmeng Wang , Ivor Tsang , Guang Dai

Image-to-image translation is a technique that focuses on transferring images from one domain to another while maintaining the essential content representations. In recent years, image-to-image translation has gained significant attention…

Image and Video Processing · Electrical Eng. & Systems 2024-04-02 Xixian Wu , Dian Chao , Yang Yang

We propose an image-to-image translation framework for facial attribute editing with disentangled interpretable latent directions. Facial attribute editing task faces the challenges of targeted attribute editing with controllable strength…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Yusuf Dalva , Hamza Pehlivan , Cansu Moran , Öykü Irmak Hatipoğlu , Ayşegül Dündar

Many real-world datasets can be divided into groups according to certain salient features (e.g. grouping images by subject, grouping text by font, etc.). Often, machine learning tasks require that these features be represented separately…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-02-16 Dan Andrei Iliescu , Aliaksei Mikhailiuk , Damon Wischik , Rafal Mantiuk

Traditionally, style has been primarily considered in terms of artistic elements such as colors, brushstrokes, and lighting. However, identical semantic subjects, like people, boats, and houses, can vary significantly across different…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Jinghao Hu , Yuhe Zhang , GuoHua Geng , Liuyuxin Yang , JiaRui Yan , Jingtao Cheng , YaDong Zhang , Kang Li

Disentangling content and style information of an image has played an important role in recent success in image translation. In this setting, how to inject given style into an input image containing its own content is an important issue,…

Computer Vision and Pattern Recognition · Computer Science 2019-12-02 Wonwoong Cho , Kangyeol Kim , Eungyeup Kim , Hyunwoo J. Kim , Jaegul Choo

Large-scale text-to-image models pre-trained on massive text-image pairs show excellent performance in image synthesis recently. However, image can provide more intuitive visual concepts than plain text. People may ask: how can we integrate…

Computer Vision and Pattern Recognition · Computer Science 2023-09-21 Bin Cheng , Zuhao Liu , Yunbo Peng , Yue Lin

This paper presents a novel design of neural network system for fine-grained style modeling, transfer and prediction in expressive text-to-speech (TTS) synthesis. Fine-grained modeling is realized by extracting style embeddings from the…

Audio and Speech Processing · Electrical Eng. & Systems 2021-10-11 Daxin Tan , Tan Lee

Recent studies show strong generative performance in domain translation especially by using transfer learning techniques on the unconditional generator. However, the control between different domain features using a single model is still…

Computer Vision and Pattern Recognition · Computer Science 2023-03-22 Dongyeun Lee , Jae Young Lee , Doyeon Kim , Jaehyun Choi , Jaejun Yoo , Junmo Kim

We consider the generic deep image enhancement problem where an input image is transformed into a perceptually better-looking image. Recent methods for image enhancement consider the problem by performing style transfer and image…

Computer Vision and Pattern Recognition · Computer Science 2020-12-14 Indra Deep Mastan , Shanmuganathan Raman

Large Vision-Language Models (LVLMs) achieve strong performance on single-image tasks, but their performance declines when multiple images are provided as input. One major reason is the cross-image information leakage, where the model…

Computer Vision and Pattern Recognition · Computer Science 2026-02-26 Minyoung Lee , Yeji Park , Dongjun Hwang , Yejin Kim , Seong Joon Oh , Junsuk Choe

Recent advances in large-scale text-to-image generation models have led to a surge in subject-driven text-to-image generation, which aims to produce customized images that align with textual descriptions while preserving the identity of…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Kewen Chen , Xiaobin Hu , Wenqi Ren

Most image-to-image translation methods focus on learning mappings across domains with the assumption that images share content (e.g., pose) but have their own domain-specific information known as style. When conditioned on a target image,…

Computer Vision and Pattern Recognition · Computer Science 2021-07-26 Mohamed Abderrahmen Abid , Ihsen Hedhli , Jean-François Lalonde , Christian Gagne

Customized text-to-image generation, which aims to learn user-specified concepts with a few images, has drawn significant attention recently. However, existing methods usually suffer from overfitting issues and entangle the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Yufei Cai , Yuxiang Wei , Zhilong Ji , Jinfeng Bai , Hu Han , Wangmeng Zuo

Training data is at the core of any successful text-to-image models. The quality and descriptiveness of image text are crucial to a model's performance. Given the noisiness and inconsistency in web-scraped datasets, recent works shifted…

Computer Vision and Pattern Recognition · Computer Science 2025-06-23 Manuel Brack , Sudeep Katakol , Felix Friedrich , Patrick Schramowski , Hareesh Ravi , Kristian Kersting , Ajinkya Kale

In this paper, we introduce a model designed to improve the prediction of image-text alignment, targeting the challenge of compositional understanding in current visual-language models. Our approach focuses on generating high-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-10-02 Yuheng Li , Haotian Liu , Mu Cai , Yijun Li , Eli Shechtman , Zhe Lin , Yong Jae Lee , Krishna Kumar Singh
‹ Prev 1 3 4 5 6 7 10 Next ›