English
Related papers

Related papers: X&Fuse: Fusing Visual Information in Text-to-Image…

200 papers

Recently, automatic image caption generation has been an important focus of the work on multimodal translation task. Existing approaches can be roughly categorized into two classes, i.e., top-down and bottom-up, the former transfers the…

Computer Vision and Pattern Recognition · Computer Science 2019-09-06 Wei Wei , Ling Cheng , Xianling Mao , Guangyou Zhou , Feida Zhu

Visual diffusion models achieve remarkable progress, yet they are typically trained at limited resolutions due to the lack of high-resolution data and constrained computation resources, hampering their ability to generate high-fidelity…

Computer Vision and Pattern Recognition · Computer Science 2025-07-14 Haonan Qiu , Shiwei Zhang , Yujie Wei , Ruihang Chu , Hangjie Yuan , Xiang Wang , Yingya Zhang , Ziwei Liu

Recent works have successfully extended large-scale text-to-image models to the video domain, producing promising results but at a high computational cost and requiring a large amount of video data. In this work, we introduce…

Computer Vision and Pattern Recognition · Computer Science 2024-05-24 Bo Peng , Xinyuan Chen , Yaohui Wang , Chaochao Lu , Yu Qiao

The aim of image captioning is to generate captions by machine to describe image contents. Despite many efforts, generating discriminative captions for images remains non-trivial. Most traditional approaches imitate the language structure…

Computer Vision and Pattern Recognition · Computer Science 2018-07-24 Xihui Liu , Hongsheng Li , Jing Shao , Dapeng Chen , Xiaogang Wang

Visual understanding is inherently contextual -- what we focus on in an image depends on the task at hand. For instance, given an image of a person holding a bouquet of flowers, we may focus on either the person such as their clothing, or…

Computer Vision and Pattern Recognition · Computer Science 2025-04-14 Cheng-Yu Hsieh , Pavan Kumar Anasosalu Vasu , Fartash Faghri , Raviteja Vemulapalli , Chun-Liang Li , Ranjay Krishna , Oncel Tuzel , Hadi Pouransari

Multi-focus image fusion is a challenging field of study that aims to provide a completely focused image by integrating focused and un-focused pixels. Most existing methods suffer from shift variance, misregistered images, and…

Computer Vision and Pattern Recognition · Computer Science 2022-04-22 Sultan Sevgi Turgut , Mustafa Oral

Information fusion is used widely to improve document classification by the integration of multiple data sources (multimodal) or representations (multiview). However, the field lacks a unified framework, a quantitative synthesis of its…

Computation and Language · Computer Science 2026-05-27 Marcin Michał Mirończuk

Current image captioning works usually focus on generating descriptions in an autoregressive manner. However, there are limited works that focus on generating descriptions non-autoregressively, which brings more decoding diversity. Inspired…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Yufeng He , Zefan Cai , Xu Gan , Baobao Chang

Large-scale text-to-image generative models have shown remarkable ability to synthesize diverse and high-quality images. However, it is still challenging to directly apply these models for editing real images for two reasons. First, it is…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Syed Muhmmad Israr , Feng Zhao

Recently, zero-shot image captioning has gained increasing attention, where only text data is available for training. The remarkable progress in text-to-image diffusion model presents the potential to resolve this task by employing…

Computer Vision and Pattern Recognition · Computer Science 2025-01-03 Jianjie Luo , Jingwen Chen , Yehao Li , Yingwei Pan , Jianlin Feng , Hongyang Chao , Ting Yao

The widespread use of large language models has resulted in a multitude of tokenizers and embedding spaces, making knowledge transfer in prompt discovery tasks difficult. In this work, we propose FUSE (Flexible Unification of Semantic…

Computation and Language · Computer Science 2024-08-12 Joshua Nathaniel Williams , J. Zico Kolter

Recently, diffusion-based image generation methods are credited for their remarkable text-to-image generation capabilities, while still facing challenges in accurately generating multilingual scene text images. To tackle this problem, we…

Computer Vision and Pattern Recognition · Computer Science 2023-12-20 Lingjun Zhang , Xinyuan Chen , Yaohui Wang , Yue Lu , Yu Qiao

Content-Based Image Retrieval (CIR) aims to search for a target image by concurrently comprehending the composition of an example image and a complementary text, which potentially impacts a wide variety of real-world applications, such as…

Artificial Intelligence · Computer Science 2022-07-12 Wenqiao Zhang , Jiannan Guo , Mengze Li , Haochen Shi , Shengyu Zhang , Juncheng Li , Siliang Tang , Yueting Zhuang

This study introduces a novel multimodal food recognition framework that effectively combines visual and textual modalities to enhance classification accuracy and robustness. The proposed approach employs a dynamic multimodal fusion…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Prateek Mittal , Puneet Goyal , Joohi Chauhan

Image reconstruction from noisy sensor measurements is challenging and many methods have been proposed for it. Yet, most approaches focus on learning robust natural image priors while modeling the scene's noise statistics. In extremely…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Erez Yosef , Raja Giryes

Personalized text-to-image generation has emerged as a powerful and sought-after tool, empowering users to create customized images based on their specific concepts and prompts. However, existing approaches to personalization encounter…

Computer Vision and Pattern Recognition · Computer Science 2023-09-13 Li Chen , Mengyi Zhao , Yiheng Liu , Mingxu Ding , Yangyang Song , Shizun Wang , Xu Wang , Hao Yang , Jing Liu , Kang Du , Min Zheng

Evaluating image captions typically relies on reference captions, which are costly to obtain and exhibit significant diversity and subjectivity. While reference-free evaluation metrics have been proposed, most focus on cross-modal…

Computer Vision and Pattern Recognition · Computer Science 2025-01-09 Tianyu Cui , Jinbin Bai , Guo-Hua Wang , Qing-Guo Chen , Zhao Xu , Weihua Luo , Kaifu Zhang , Ye Shi

In recent years, there has been significant progress in the development of text-to-image generative models. Evaluating the quality of the generative models is one essential step in the development process. Unfortunately, the evaluation…

Computer Vision and Pattern Recognition · Computer Science 2024-03-26 Lin Zhao , Tianchen Zhao , Zinan Lin , Xuefei Ning , Guohao Dai , Huazhong Yang , Yu Wang

Generative diffusion models offer a natural choice for data augmentation when training complex vision models. However, ensuring reliability of their generative content as augmentation samples remains an open challenge. Despite a number of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Khawar Islam , Naveed Akhtar

This dissertation attempts to drive innovation in the field of generative modeling for computer vision, by exploring novel formulations of conditional generative models, and innovative applications in images, 3D animations, and video. Our…

Computer Vision and Pattern Recognition · Computer Science 2023-10-23 Vikram Voleti