English
Related papers

Related papers: Text-to-Image Diffusion Models are Great Sketch-Ph…

200 papers

Large-scale text-to-image generative models have shown their remarkable ability to synthesize diverse and high-quality images. However, it is still challenging to directly apply these models for editing real images for two reasons. First,…

Computer Vision and Pattern Recognition · Computer Science 2023-02-07 Gaurav Parmar , Krishna Kumar Singh , Richard Zhang , Yijun Li , Jingwan Lu , Jun-Yan Zhu

This paper investigates the use of large-scale diffusion models for Zero-Shot Video Object Segmentation (ZS-VOS) without fine-tuning on video data or training on any image segmentation data. While diffusion models have demonstrated strong…

Computer Vision and Pattern Recognition · Computer Science 2025-04-09 Thanos Delatolas , Vicky Kalogeiton , Dim P. Papadopoulos

The recent wave of large-scale text-to-image diffusion models has dramatically increased our text-based image generation abilities. These models can generate realistic images for a staggering variety of prompts and exhibit impressive…

Machine Learning · Computer Science 2023-09-14 Alexander C. Li , Mihir Prabhudesai , Shivam Duggal , Ellis Brown , Deepak Pathak

Most existing algorithms for cross-modal Information Retrieval are based on a supervised train-test setup, where a model learns to align the mode of the query (e.g., text) to the mode of the documents (e.g., images) from a given training…

Computer Vision and Pattern Recognition · Computer Science 2020-09-24 Anurag Roy , Vinay Kumar Verma , Kripabandhu Ghosh , Saptarshi Ghosh

Text-to-Image models have introduced a remarkable leap in the evolution of machine learning, demonstrating high-quality synthesis of images from a given text-prompt. However, these powerful pretrained models still lack control handles that…

Computer Vision and Pattern Recognition · Computer Science 2022-11-28 Andrey Voynov , Kfir Aberman , Daniel Cohen-Or

Learning from feedback has been shown to enhance the alignment between text prompts and images in text-to-image diffusion models. However, due to the lack of focus in feedback content, especially regarding the object type and quantity,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-03 Xuexiang Niu , Jinping Tang , Lei Wang , Ge Zhu

Zero-shot sketch-based image retrieval (ZS-SBIR) is a specific cross-modal retrieval task for searching natural images given free-hand sketches under the zero-shot scenario. Most existing methods solve this problem by simultaneously…

Computer Vision and Pattern Recognition · Computer Science 2022-05-09 Xinxun Xu , Muli Yang , Yanhua Yang , Hao Wang

Unsupervised visual object tracking is a challenging task that requires following arbitrary targets in videos without training on ground-truth annotations. Despite considerable progress, existing state-of-the-art unsupervised trackers often…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Zhengbo Zhang , Zhigang Tu , Junsong Yuan , De Wen Soh , Bo Du

Traditionally, style has been primarily considered in terms of artistic elements such as colors, brushstrokes, and lighting. However, identical semantic subjects, like people, boats, and houses, can vary significantly across different…

Computer Vision and Pattern Recognition · Computer Science 2024-10-25 Jinghao Hu , Yuhe Zhang , GuoHua Geng , Liuyuxin Yang , JiaRui Yan , Jingtao Cheng , YaDong Zhang , Kang Li

Diffusion models, such as Stable Diffusion, have shown incredible performance on text-to-image generation. Since text-to-image generation often requires models to generate visual concepts with fine-grained details and attributes specified…

Computer Vision and Pattern Recognition · Computer Science 2024-04-26 Xuehai He , Weixi Feng , Tsu-Jui Fu , Varun Jampani , Arjun Akula , Pradyumna Narayana , Sugato Basu , William Yang Wang , Xin Eric Wang

This paper studies the problem of zero-shot sketch-based image retrieval (ZS-SBIR), which aims to use sketches from unseen categories as queries to match the images of the same category. Due to the large cross-modality discrepancy, ZS-SBIR…

Computer Vision and Pattern Recognition · Computer Science 2023-12-19 Decheng Liu , Xu Luo , Chunlei Peng , Nannan Wang , Ruimin Hu , Xinbo Gao

Recent advances in text-to-image diffusion models have substantially improved the quality of image customization, enabling the synthesis of highly realistic images. Despite this progress, achieving fast and efficient personalization remains…

Computer Vision and Pattern Recognition · Computer Science 2026-03-25 Aniket Roy , Maitreya Suin , Rama Chellappa

The problem of zero-shot sketch-based image retrieval (ZS-SBIR) has achieved increasing attention due to its wide applications, e.g. e-commerce. Despite progress made in this field, previous works suffer from using imbalanced samples of…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Hanwen Su , Ge Song , Jiyan Wang , Yuanbo Zhu

Fine-grained image retrieval via hand-drawn sketches or textual descriptions remains a critical challenge due to inherent modality gaps. While hand-drawn sketches capture complex structural contours, they lack color and texture, which text…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Siyuan Wang , Hanchen Gao , Guangming Zhu , Jiang Lu , Yiyue Ma , Tianci Wu , Jincai Huang , Liang Zhang

Diffusion models have dramatically advanced text-to-image generation in recent years, translating abstract concepts into high-fidelity images with remarkable ease. In this work, we examine whether they can also blend distinct concepts,…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Lorenzo Olearo , Giorgio Longari , Alessandro Raganato , Rafael Peñaloza , Simone Melzi

Sketch-based image retrieval (SBIR) is the task of retrieving images from a natural image database that correspond to a given hand-drawn sketch. Ideally, an SBIR model should learn to associate components in the sketch (say, feet, tail,…

Computer Vision and Pattern Recognition · Computer Science 2018-08-01 Sasi Kiran Yelamarthi , Shiva Krishna Reddy , Ashish Mishra , Anurag Mittal

The efficacy of zero-shot sketch-based image retrieval (ZS-SBIR) models is governed by two challenges. The immense distributions-gap between the sketches and the images requires a proper domain alignment. Moreover, the fine-grained nature…

Computer Vision and Pattern Recognition · Computer Science 2022-01-19 Ushasi Chaudhuri , Ruchika Chavan , Biplab Banerjee , Anjan Dutta , Zeynep Akata

Diffusion models have demonstrated impressive performance in text-guided image generation. Current methods that leverage the knowledge of these models for image editing either fine-tune them using the input image (e.g., Imagic) or…

Computer Vision and Pattern Recognition · Computer Science 2023-11-09 Zhongping Zhang , Jian Zheng , Jacob Zhiyuan Fang , Bryan A. Plummer

Restoring low-resolution text images presents a significant challenge, as it requires maintaining both the fidelity and stylistic realism of the text in restored images. Existing text image restoration methods often fall short in hard…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Chenglu Pan , Xiaogang Xu , Ganggui Ding , Yunke Zhang , Wenbo Li , Jiarong Xu , Qingbiao Wu

The text-to-image synthesis by diffusion models has recently shown remarkable performance in generating high-quality images. Although performs well for simple texts, the models may get confused when faced with complex texts that contain…

Computer Vision and Pattern Recognition · Computer Science 2024-01-15 Chang Yu , Junran Peng , Xiangyu Zhu , Zhaoxiang Zhang , Qi Tian , Zhen Lei