English
Related papers

Related papers: ShareGPT-4o-Image: Aligning Multimodal Models with…

200 papers

Unsupervised image-to-image translation aims to learn the translation between two visual domains without paired data. Despite the recent progress in image translation models, it remains challenging to build mappings between complex domains…

Computer Vision and Pattern Recognition · Computer Science 2022-04-08 Shuai Yang , Liming Jiang , Ziwei Liu , Chen Change Loy

We provide a dataset for enabling Deep Generative Models (DGMs) in engineering design and propose methods to automate data labeling by utilizing large-scale foundation models. GeoBiked is curated to contain 4 355 bicycle images, annotated…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Phillip Mueller , Sebastian Mueller , Lars Mikelsons

In this paper, we critically evaluate the capabilities of the state-of-the-art multimodal large language model, i.e., GPT-4 with Vision (GPT-4V), on Visual Question Answering (VQA) task. Our experiments thoroughly assess GPT-4V's…

Computer Vision and Pattern Recognition · Computer Science 2023-10-31 Zhiling Yan , Kai Zhang , Rong Zhou , Lifang He , Xiang Li , Lichao Sun

Generative AI based on foundation models provides a first glimpse into the world represented by machines trained on vast amounts of multimodal data ingested by these models during training. If we consider the resulting models as knowledge…

Computers and Society · Computer Science 2024-04-12 Zilong Liu , Krzysztof Janowicz , Kitty Currier , Meilin Shi

Text-to-image (T2I) models have recently experienced rapid development, achieving astonishing performance in terms of fidelity and textual alignment capabilities. However, given a long paragraph (up to 512 words), these generation models…

Computer Vision and Pattern Recognition · Computer Science 2025-05-07 Weijia Wu , Zhuang Li , Yefei He , Mike Zheng Shou , Chunhua Shen , Lele Cheng , Yan Li , Tingting Gao , Di Zhang

This work conducts an evaluation of GPT-4V's multimodal capability for medical image analysis, with a focus on three representative tasks of radiology report generation, medical visual question answering, and medical visual grounding. For…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Yingshu Li , Yunyi Liu , Zhanyu Wang , Xinyu Liang , Lei Wang , Lingqiao Liu , Leyang Cui , Zhaopeng Tu , Longyue Wang , Luping Zhou

Existing text-to-image diffusion models primarily generate images from text prompts. However, the inherent conciseness of textual descriptions poses challenges in faithfully synthesizing images with intricate details, such as specific…

Computer Vision and Pattern Recognition · Computer Science 2024-06-07 Wei Li , Xue Xu , Jiachen Liu , Xinyan Xiao

We propose to use automatically generated instruction-following data to improve the zero-shot capabilities of a large multimodal model with additional support for generative and image editing tasks. We achieve this by curating a new…

Computer Vision and Pattern Recognition · Computer Science 2024-10-04 Jefferson Hernandez , Ruben Villegas , Vicente Ordonez

We introduce $\textbf{Ovis-Image}$, a 7B text-to-image model specifically optimized for high-quality text rendering, designed to operate efficiently under stringent computational constraints. Built upon our previous Ovis-U1 framework,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-01 Guo-Hua Wang , Liangfu Cao , Tianyu Cui , Minghao Fu , Xiaohao Chen , Pengxin Zhan , Jianshan Zhao , Lan Li , Bowen Fu , Jiaqi Liu , Qing-Guo Chen

The study evaluates and compares GPT-4 and GPT-4Vision for radiological tasks, suggesting GPT-4Vision may recognize radiological features from images, thereby enhancing its diagnostic potential over text-based descriptions.

Image and Video Processing · Electrical Eng. & Systems 2026-01-23 Felix Busch , Tianyu Han , Marcus Makowski , Daniel Truhn , Keno Bressem , Lisa Adams

This paper presents an energy-efficient stable diffusion processor for text-to-image generation. While stable diffusion attained attention for high-quality image synthesis results, its inherent characteristics hinder its deployment on…

Hardware Architecture · Computer Science 2024-09-24 Jiwon Choi , Wooyoung Jo , Seongyon Hong , Beomseok Kwon , Wonhoon Park , Hoi-Jun Yoo

We present 4DNeX, the first feed-forward framework for generating 4D (i.e., dynamic 3D) scene representations from a single image. In contrast to existing methods that rely on computationally intensive optimization or require multi-frame…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Zhaoxi Chen , Tianqi Liu , Long Zhuo , Jiawei Ren , Zeng Tao , He Zhu , Fangzhou Hong , Liang Pan , Ziwei Liu

Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if they are two separate functions encapsulated within the same…

Computer Vision and Pattern Recognition · Computer Science 2025-10-22 Kaihang Pan , Yang Wu , Wendong Bu , Kai Shen , Juncheng Li , Yingting Wang , Yunfei Li , Siliang Tang , Jun Xiao , Fei Wu , Hang Zhao , Yueting Zhuang

We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at…

We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs -- images, text, audio, action,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Jiasen Lu , Christopher Clark , Sangho Lee , Zichen Zhang , Savya Khosla , Ryan Marten , Derek Hoiem , Aniruddha Kembhavi

Recent developments in multimodal large language models (MLLMs) have spurred significant interest in their potential applications across various medical imaging domains. On the one hand, there is a temptation to use these generative models…

Image and Video Processing · Electrical Eng. & Systems 2024-06-05 Sulaiman Khan , Md. Rafiul Biswas , Alina Murad , Hazrat Ali , Zubair Shah

Due to the fascinating generative performance of text-to-image diffusion models, growing text-to-3D generation works explore distilling the 2D generative priors into 3D, using the score distillation sampling (SDS) loss, to bypass the data…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Yu-Jie Yuan , Leif Kobbelt , Jiwen Liu , Yuan Zhang , Pengfei Wan , Yu-Kun Lai , Lin Gao

Despite that the performance of image-to-image translation has been significantly improved by recent progress in generative models, current methods still suffer from severe degradation in training stability and sample quality when applied…

Computer Vision and Pattern Recognition · Computer Science 2019-04-16 Jie Cao , Huaibo Huang , Yi Li , Jingtuo Liu , Ran He , Zhenan Sun

Diffusion models have achieved remarkable advancements in text-to-image generation. However, existing models still have many difficulties when faced with multiple-object compositional generation. In this paper, we propose RealCompo, a new…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Xinchen Zhang , Ling Yang , Yaqi Cai , Zhaochen Yu , Kai-Ni Wang , Jiake Xie , Ye Tian , Minkai Xu , Yong Tang , Yujiu Yang , Bin Cui

Advances in generative artificial intelligence have altered multimedia creation, allowing for automatic cinematic video synthesis from text inputs. This work describes a method for creating 60-second cinematic movies incorporating Stable…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Sridhar S , Nithin A , Shakeel Rifath , Vasantha Raj