English
Related papers

Related papers: MUMU: Bootstrapping Multimodal Image Generation fr…

200 papers

We introduce MarkupDM, a multimodal markup document model that represents graphic design as an interleaved multimodal document consisting of both markup language and images. Unlike existing holistic approaches that rely on an…

Computer Vision and Pattern Recognition · Computer Science 2025-12-05 Kotaro Kikuchi , Ukyo Honda , Naoto Inoue , Mayu Otani , Edgar Simo-Serra , Kota Yamaguchi

Image captioning is a fast-growing research field of computer vision and natural language processing that involves creating text explanations for images. This study aims to develop a system that uses a pre-trained convolutional neural…

Computation and Language · Computer Science 2022-03-04 Rashid Khan , M Shujah Islam , Khadija Kanwal , Mansoor Iqbal , Md. Imran Hossain , Zhongfu Ye

Recently, the diffusion-based generative paradigm has achieved impressive general image generation capabilities with text prompts due to its accurate distribution modeling and stable training process. However, generating diverse remote…

Image and Video Processing · Electrical Eng. & Systems 2024-10-31 Jialin Luo , Yuanzhi Wang , Ziqi Gu , Yide Qiu , Shuaizhen Yao , Fuyun Wang , Chunyan Xu , Wenhua Zhang , Dan Wang , Zhen Cui

Teaching robots novel behaviors typically requires motion demonstrations via teleoperation or kinaesthetic teaching, that is, physically guiding the robot. While recent work has explored using human sketches to specify desired behaviors,…

Robotics · Computer Science 2025-09-26 William Barron , Xiaoxiang Dong , Matthew Johnson-Roberson , Weiming Zhi

Unified multimodal models (UMMs) aim to integrate multimodal understanding and generation within a unified architecture, yet it remains unclear to what extent their representations are truly aligned across modalities. To investigate this…

Computation and Language · Computer Science 2026-04-08 Cheng Yang , Chufan Shi , Bo Shui , Yaokang Wu , Muzi Tao , Huijuan Wang , Ivan Yee Lee , Yong Liu , Xuezhe Ma , Taylor Berg-Kirkpatrick

Unified multimodal models aim to jointly enable visual understanding and generation, yet current benchmarks rarely examine their true integration. Existing evaluations either treat the two abilities in isolation or overlook tasks that…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Kai Zou , Ziqi Huang , Yuhao Dong , Shulin Tian , Dian Zheng , Hongbo Liu , Jingwen He , Bin Liu , Yu Qiao , Ziwei Liu

Image captioning has demonstrated models that are capable of generating plausible text given input images or videos. Further, recent work in image generation has shown significant improvements in image quality when text is used as a prior.…

Machine Learning · Computer Science 2018-09-28 Shagan Sah , Dheeraj Peri , Ameya Shringi , Chi Zhang , Miguel Dominguez , Andreas Savakis , Ray Ptucha

To develop high-performing Visual Language Models (VLMs), it is essential to prepare multimodal resources, such as image-text pairs, interleaved data, and instruction data. While multimodal resources for English are abundant, there is a…

Computation and Language · Computer Science 2024-10-31 Keito Sasagawa , Koki Maeda , Issa Sugiura , Shuhei Kurita , Naoaki Okazaki , Daisuke Kawahara

The multifaceted nature of human perception and comprehension indicates that, when we think, our body can naturally take any combination of senses, a.k.a., modalities and form a beautiful picture in our brain. For example, when we see a…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Yuanhuiyi Lyu , Xu Zheng , Lin Wang

We review research on generating visual data from text from the angle of "cross-modal generation." This point of view allows us to draw parallels between various methods geared towards working on input text and producing visual output,…

Computer Vision and Pattern Recognition · Computer Science 2024-01-23 Maciej Żelaszczyk , Jacek Mańdziuk

Instruction-based image editing holds immense potential for a variety of applications, as it enables users to perform any editing operation using a natural language instruction. However, current models in this domain often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2023-11-17 Shelly Sheynin , Adam Polyak , Uriel Singer , Yuval Kirstain , Amit Zohar , Oron Ashual , Devi Parikh , Yaniv Taigman

Combining the visual modality with pretrained language models has been surprisingly effective for simple descriptive tasks such as image captioning. More general text generation however remains elusive. We take a step back and ask: How do…

Computation and Language · Computer Science 2022-10-25 Shruti Palaskar , Akshita Bhagia , Yonatan Bisk , Florian Metze , Alan W Black , Ana Marasović

Recent advancements in instruction-based image editing and subject-driven generation have garnered significant attention, yet both tasks still face limitations in meeting practical user needs. Instruction-based editing relies solely on…

Computer Vision and Pattern Recognition · Computer Science 2025-10-09 Bin Xia , Bohao Peng , Yuechen Zhang , Junjia Huang , Jiyang Liu , Jingyao Li , Haoru Tan , Sitong Wu , Chengyao Wang , Yitong Wang , Xinglong Wu , Bei Yu , Jiaya Jia

We extend the SKIP-GRAM model of Mikolov et al. (2013a) by taking visual information into account. Like SKIP-GRAM, our multimodal models (MMSKIP-GRAM) build vector-based word representations by learning to predict linguistic contexts in…

Computation and Language · Computer Science 2015-03-13 Angeliki Lazaridou , Nghia The Pham , Marco Baroni

Prior methods for controlling image generation are limited in their ability to be taught new tasks. In contrast, vision-language models, or VLMs, can learn tasks in-context and produce the correct outputs for a given input. We propose a…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Grace Luo , Jonathan Granskog , Aleksander Holynski , Trevor Darrell

This work aims to create a multimodal AI system that chats with humans and shares relevant photos. While earlier works were limited to dialogues about specific objects or scenes within images, recent works have incorporated images into…

Computation and Language · Computer Science 2023-05-08 Min Young Lee

Text-conditioned image generation models are a prevalent use of AI image synthesis, yet intuitively controlling output guided by an artist remains challenging. Current methods require multiple images and textual prompts for each object to…

Computer Vision and Pattern Recognition · Computer Science 2024-01-02 Shounak Chatterjee

Recent multimodal large language models (MLLMs) have shown promising instruction following capabilities on vision-language tasks. In this work, we introduce VISUAL MODALITY INSTRUCTION (VIM), and investigate how well multimodal models can…

Computer Vision and Pattern Recognition · Computer Science 2024-06-12 Xiujun Li , Yujie Lu , Zhe Gan , Jianfeng Gao , William Yang Wang , Yejin Choi

We consider the problem of generating free-form mobile manipulation instructions based on a target object image and receptacle image. Conventional image captioning models are not able to generate appropriate instructions because their…

Robotics · Computer Science 2025-01-29 Kei Katsumata , Motonari Kambara , Daichi Yashima , Ryosuke Korekata , Komei Sugiura

We present Unified-IO 2, the first autoregressive multimodal model that is capable of understanding and generating image, text, audio, and action. To unify different modalities, we tokenize inputs and outputs -- images, text, audio, action,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-29 Jiasen Lu , Christopher Clark , Sangho Lee , Zichen Zhang , Savya Khosla , Ryan Marten , Derek Hoiem , Aniruddha Kembhavi