English
Related papers

Related papers: ViSTA: Visual Storytelling using Multi-modal Adapt…

200 papers

Visual storytelling is an interdisciplinary field combining computer vision and natural language processing to generate cohesive narratives from sequences of images. This paper presents a novel approach that leverages recent advancements in…

Computation and Language · Computer Science 2025-06-11 Mohamed Gado , Towhid Taliee , Muhammad Memon , Dmitry Ignatov , Radu Timofte

Text-to-image diffusion models rely on text embeddings from a pre-trained text encoder, but these embeddings remain fixed across all diffusion timesteps, limiting their adaptability to the generative process. We propose Diffusion Adaptive…

Machine Learning · Computer Science 2025-10-29 Byeonghu Na , Minsang Park , Gyuwon Sim , Donghyeok Shin , HeeSun Bae , Mina Kang , Se Jung Kwon , Wanmo Kang , Il-Chul Moon

This paper introduces ITA-MDT, the Image-Timestep-Adaptive Masked Diffusion Transformer Framework for Image-Based Virtual Try-On (IVTON), designed to overcome the limitations of previous approaches by leveraging the Masked Diffusion…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Ji Woo Hong , Tri Ton , Trung X. Pham , Gwanhyeong Koo , Sunjae Yoon , Chang D. Yoo

Information extraction, e.g., attribute value extraction, has been extensively studied and formulated based only on text. However, many attributes can benefit from image-based extraction, like color, shape, pattern, among others. The visual…

Computation and Language · Computer Science 2023-06-05 Hejie Cui , Rongmei Lin , Nasser Zalmout , Chenwei Zhang , Jingbo Shang , Carl Yang , Xian Li

Existing multi-modal image fusion methods fail to address the compound degradations presented in source images, resulting in fusion images plagued by noise, color bias, improper exposure, \textit{etc}. Additionally, these methods often…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Hao Zhang , Lei Cao , Jiayi Ma

Adapting models to dynamic, real-world environments characterized by shifting data distributions and unseen test scenarios is a critical challenge in deep learning. In this paper, we consider a realistic and challenging Test-Time Adaptation…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Manogna Sreenivas , Soma Biswas

Recent advances in generative diffusion models have enabled text-controlled synthesis of realistic and diverse images with impressive quality. Despite these remarkable advances, the application of text-to-image generative models in computer…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Yulu Gan , Sungwoo Park , Alexander Schubert , Anthony Philippakis , Ahmed M. Alaa

Image animation has seen significant progress, driven by the powerful generative capabilities of diffusion models. However, maintaining appearance consistency with static input images and mitigating abrupt motion transitions in generated…

Computer Vision and Pattern Recognition · Computer Science 2025-08-12 Xin Ma , Yaohui Wang , Genyun Jia , Xinyuan Chen , Tien-Tsin Wong , Cunjian Chen

Image-to-image translation aims to learn a mapping between a source and a target domain, enabling tasks such as style transfer, appearance transformation, and domain adaptation. In this work, we explore a diffusion-based framework for…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Qiang Zhu , Kuan Lu , Menghao Huo , Yuxiao Li

Diffusion-based text-to-image models have rapidly gained popularity for their ability to generate detailed and realistic images from textual descriptions. However, these models often reflect the biases present in their training data,…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Hidir Yesiltepe , Kiymet Akdemir , Pinar Yanardag

Text-to-image diffusion models have demonstrated tremendous success in synthesizing visually stunning images given textual instructions. Despite remarkable progress in creating high-fidelity visuals, text-to-image models can still struggle…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Taewook Kim , Ze Wang , Zhengyuan Yang , Jiang Wang , Lijuan Wang , Zicheng Liu , Qiang Qiu

The goal of this paper is to optimize the training process of diffusion-based text-to-speech models. While recent studies have achieved remarkable advancements, their training demands substantial time and computational costs, largely due to…

Audio and Speech Processing · Electrical Eng. & Systems 2025-06-02 Jeongsoo Choi , Zhikang Niu , Ji-Hoon Kim , Chunhui Wang , Joon Son Chung , Xie Chen

With the rapid progression of deep learning technologies, multi-modality image fusion has become increasingly prevalent in object detection tasks. Despite its popularity, the inherent disparities in how different sources depict scene…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Xingyuan Li , Yang Zou , Jinyuan Liu , Zhiying Jiang , Long Ma , Xin Fan , Risheng Liu

Story visualization aims to generate a series of realistic and coherent images based on a storyline. Current models adopt a frame-by-frame architecture by transforming the pre-trained text-to-image model into an auto-regressive manner.…

Computer Vision and Pattern Recognition · Computer Science 2024-04-10 Ming Tao , Bing-Kun Bao , Hao Tang , Yaowei Wang , Changsheng Xu

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

Sound · Computer Science 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

In this paper, we address the limitations of existing text-to-image diffusion models in generating demographically fair results when given human-related descriptions. These models often struggle to disentangle the target language context…

Computer Vision and Pattern Recognition · Computer Science 2024-03-07 Jia Li , Lijie Hu , Jingfeng Zhang , Tianhang Zheng , Hua Zhang , Di Wang

Diffusion models have gained tremendous success in text-to-image generation, yet still lag behind with visual understanding tasks, an area dominated by autoregressive vision-language models. We propose a large-scale and fully end-to-end…

Computer Vision and Pattern Recognition · Computer Science 2025-04-03 Zijie Li , Henry Li , Yichun Shi , Amir Barati Farimani , Yuval Kluger , Linjie Yang , Peng Wang

We present a simple but effective training-free approach for text-driven image-to-image translation based on a pretrained text-to-image diffusion model. Our goal is to generate an image that aligns with the target task while preserving the…

Computer Vision and Pattern Recognition · Computer Science 2024-12-23 Hyunsoo Lee , Minsoo Kang , Bohyung Han

As artificial intelligence-generated content (AIGC) continues to evolve, video-to-audio (V2A) generation has emerged as a key area with promising applications in multimedia editing, augmented reality, and automated content creation. While…

Computer Vision and Pattern Recognition · Computer Science 2025-03-17 Yuhuan You , Xihong Wu , Tianshu Qu

Incorporating auxiliary modalities such as images into event detection models has attracted increasing interest over the last few years. The complexity of natural language in describing situations has motivated researchers to leverage the…

Computation and Language · Computer Science 2023-06-06 Farhad Moghimifar , Fatemeh Shiri , Van Nguyen , Reza Haffari , Yuan-Fang Li