English
Related papers

Related papers: Adaptive Visual Conditioning for Semantic Consiste…

200 papers

Open-vocabulary semantic segmentation requires models to effectively integrate visual representations with open-vocabulary semantic labels. While Contrastive Language-Image Pre-training (CLIP) models shine in recognizing visual concepts…

Computer Vision and Pattern Recognition · Computer Science 2024-08-12 Mengcheng Lan , Chaofeng Chen , Yiping Ke , Xinjiang Wang , Litong Feng , Wayne Zhang

Diffusion models are generative models with impressive text-to-image synthesis capabilities and have spurred a new wave of creative methods for classical machine learning tasks. However, the best way to harness the perceptual knowledge of…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Neehar Kondapaneni , Markus Marks , Manuel Knott , Rogerio Guimaraes , Pietro Perona

Semantic segmentation models trained on synthetic data often perform poorly on real-world images due to domain gaps, particularly in adverse conditions where labeled data is scarce. Yet, recent foundation models enable to generate realistic…

Computer Vision and Pattern Recognition · Computer Science 2025-09-19 Estelle Chigot , Dennis G. Wilson , Meriem Ghrib , Thomas Oberlin

Image retrieval from contextual descriptions (IRCD) aims to identify an image within a set of minimally contrastive candidates based on linguistically complex text. Despite the success of VLMs, they still significantly lag behind human…

Computer Vision and Pattern Recognition · Computer Science 2024-05-30 Honglin Lin , Siyu Li , Guoshun Nan , Chaoyue Tang , Xueting Wang , Jingxin Xu , Rong Yankai , Zhili Zhou , Yutong Gao , Qimei Cui , Xiaofeng Tao

This paper describes our zero-shot approaches for the Visual Word Sense Disambiguation (VWSD) Task in English. Our preliminary study shows that the simple approach of matching candidate images with the phrase using CLIP suffers from the…

Computation and Language · Computer Science 2023-07-13 Jie S. Li , Yow-Ting Shiue , Yong-Siang Shih , Jonas Geiping

Referring image segmentation is a challenging task that involves generating pixel-wise segmentation masks based on natural language descriptions. The complexity of this task increases with the intricacy of the sentences provided. Existing…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Hai Nguyen-Truong , E-Ro Nguyen , Tuan-Anh Vu , Minh-Triet Tran , Binh-Son Hua , Sai-Kit Yeung

Feature representation plays a crucial role in visual correspondence, and recent methods for image matching resort to deeply stacked convolutional layers. These models, however, are both monolithic and static in the sense that they…

Computer Vision and Pattern Recognition · Computer Science 2020-07-22 Juhong Min , Jongmin Lee , Jean Ponce , Minsu Cho

Diffusion models have shown significant progress in image translation tasks recently. However, due to their stochastic nature, there's often a trade-off between style transformation and content preservation. Current strategies aim to…

Computer Vision and Pattern Recognition · Computer Science 2023-06-08 Gihyun Kwon , Jong Chul Ye

Diffusion-based image editing is a composite process of preserving the source image content and generating new content or applying modifications. While current editing approaches have made improvements under text guidance, most of them have…

Computer Vision and Pattern Recognition · Computer Science 2024-03-18 Tianrui Huang , Pu Cao , Lu Yang , Chun Liu , Mengjie Hu , Zhiwei Liu , Qing Song

Visual Question Answering (VQA) is a multi-modal task that involves answering questions from an input image, semantically understanding the contents of the image and answering it in natural language. Using VQA for disaster management is an…

Computer Vision and Pattern Recognition · Computer Science 2022-11-14 Aditya Kane , V Manushree , Sahil Khose

Visual-textual correlations in the attention maps derived from text-to-image diffusion models are proven beneficial to dense visual prediction tasks, e.g., semantic segmentation. However, a significant challenge arises due to the input…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Jiayi Lin , Jiabo Huang , Jian Hu , Shaogang Gong

Latest methods for visual counterfactual explanations (VCE) harness the power of deep generative models to synthesize new examples of high-dimensional images of impressive quality. However, it is currently difficult to compare the…

Computer Vision and Pattern Recognition · Computer Science 2023-08-14 Philipp Vaeth , Alexander M. Fruehwald , Benjamin Paassen , Magda Gregorova

The task of 3D shape captioning occupies a significant place within the domain of computer graphics and has garnered considerable interest in recent years. Traditional approaches to this challenge frequently depend on the utilization of…

Graphics · Computer Science 2025-09-30 Zhenyu Shu , Jiawei Wen , Shiyang Li , Shiqing Xin , Ligang Liu

Generating a coherent sequence of images that tells a visual story, using text-to-image diffusion models, often faces the critical challenge of maintaining subject consistency across all story scenes. Existing approaches, which typically…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Gopalji Gaur , Mohammadreza Zolfaghari , Thomas Brox

We present a framework for high-fidelity product image recontextualization using text-to-image diffusion models and a novel data augmentation pipeline. This pipeline leverages image-to-video diffusion, in/outpainting & negatives to create…

Computer Vision and Pattern Recognition · Computer Science 2025-03-13 Ishaan Malhi , Praneet Dutta , Ellie Talius , Sally Ma , Brendan Driscoll , Krista Holden , Garima Pruthi , Arunachalam Narayanaswamy

The Stable Diffusion model is a prominent text-to-image generation model that relies on a text prompt as its input, which is encoded using the Contrastive Language-Image Pre-Training (CLIP). However, text prompts have limitations when it…

Computer Vision and Pattern Recognition · Computer Science 2024-02-16 Yuxuan Ding , Chunna Tian , Haoxuan Ding , Lingqiao Liu

Continual learning (CL) aims to equip models with the ability to learn from a stream of tasks without forgetting previous knowledge. With the progress of vision-language models like Contrastive Language-Image Pre-training (CLIP), their…

Computer Vision and Pattern Recognition · Computer Science 2025-11-12 Lingfeng He , De Cheng , Di Xu , Huaijie Wang , Nannan Wang

Generating images with embedded text is crucial for the automatic production of visual and multimodal documents, such as educational materials and advertisements. However, existing diffusion-based text-to-image models often struggle to…

Computer Vision and Pattern Recognition · Computer Science 2025-03-19 Forouzan Fallah , Maitreya Patel , Agneet Chatterjee , Vlad I. Morariu , Chitta Baral , Yezhou Yang

Recent work has studied text-to-audio synthesis using large amounts of paired text-audio data. However, audio recordings with high-quality text annotations can be difficult to acquire. In this work, we approach text-to-audio synthesis using…

Despite significant progress in image captioning, generating accurate and descriptive captions remains a long-standing challenge. In this study, we propose Attention-Guided Image Captioning (AGIC), which amplifies salient visual regions…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 L. D. M. S. Sai Teja , Ashok Urlana , Pruthwik Mishra