English
Related papers

Related papers: IMAGAgent: Orchestrating Multi-Turn Image Editing …

200 papers

We propose LLM-Interleaved (LLM-I), a flexible and dynamic framework that reframes interleaved image-text generation as a tool-use problem. LLM-I is designed to overcome the "one-tool" bottleneck of current unified models, which are limited…

Machine Learning · Computer Science 2025-09-18 Zirun Guo , Feng Zhang , Kai Jia , Tao Jin

Recent advances in diffusion models can generate high-quality and stunning images from text. However, multi-turn image generation, which is of high demand in real-world scenarios, still faces challenges in maintaining semantic consistency…

Computer Vision and Pattern Recognition · Computer Science 2025-06-02 Junhao Cheng , Baiqiao Yin , Kaixin Cai , Minbin Huang , Hanhui Li , Yuxin He , Xi Lu , Yue Li , Yifei Li , Yuhao Cheng , Yiqiang Yan , Xiaodan Liang

Humans solve problems by executing targeted plans, yet large language models (LLMs) remain unreliable for structured workflow execution. We propose RunAgent, a multi-agent plan execution platform that interprets natural-language plans while…

Machine Learning · Computer Science 2026-05-04 Arunabh Srivastava , Mohammad A. , Khojastepour , Srimat Chakradhar , Sennur Ulukus

Large language models (LLMs) have emerged as the dominant paradigm for robotic task planning using natural language instructions. However, trained on general internet data, LLMs are not inherently aligned with the embodiment, skill sets,…

We present OmniBooth, an image generation framework that enables spatial control with instance-level multi-modal customization. For all instances, the multimodal instruction can be described through text prompts or image references. Given a…

Computer Vision and Pattern Recognition · Computer Science 2024-10-08 Leheng Li , Weichao Qiu , Xu Yan , Jing He , Kaiqiang Zhou , Yingjie Cai , Qing Lian , Bingbing Liu , Ying-Cong Chen

Vision-language models (VLMs) like CLIP have showcased a remarkable ability to extract transferable features for downstream tasks. Nonetheless, the training process of these models is usually based on a coarse-grained contrastive loss…

Computer Vision and Pattern Recognition · Computer Science 2024-09-13 Ali Abdollah , Amirmohammad Izadi , Armin Saghafian , Reza Vahidimajd , Mohammad Mozafari , Amirreza Mirzaei , Mohammadmahdi Samiei , Mahdieh Soleymani Baghshah

In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generation tasks, including text-to-image, image editing, and in-context generation. Unlike OmniGen v1,…

Complex image restoration aims to recover high-quality images from inputs affected by multiple degradations such as blur, noise, rain, and compression artifacts. Recent restoration agents, powered by vision-language models and large…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Jianglin Lu , Yuanwei Wu , Ziyi Zhao , Hongcheng Wang , Felix Jimenez , Abrar Majeedi , Yun Fu

Large Vision-Language Models (LVLMs) have transformed multi-modal understanding, excelling in tasks like image captioning and visual question answering by integrating visual and textual inputs. However, their robustness against adversarial…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Xiang Fang , Wanlong Fang , Changshuo Wang

Recent text-to-image models produce high-quality results but still struggle with precise visual control, balancing multimodal inputs, and requiring extensive training for complex multimodal image generation. To address these limitations, we…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Haozhe Zhao , Zefan Cai , Shuzheng Si , Liang Chen , Jiuxiang Gu , Wen Xiao , Minjia Zhang , Junjie Hu

Large Language Models (LLMs) and Visual Language Models (VLMs) are attracting increasing interest due to their improving performance and applications across various domains and tasks. However, LLMs and VLMs can produce erroneous results,…

Artificial Intelligence · Computer Science 2024-12-31 Michele Brienza , Francesco Argenziano , Vincenzo Suriani , Domenico D. Bloisi , Daniele Nardi

Foundation models are becoming valuable tools in medicine. Yet despite their promise, the best way to leverage Large Language Models (LLMs) in complex medical tasks remains an open question. We introduce a novel multi-agent framework, named…

Computation and Language · Computer Science 2024-10-31 Yubin Kim , Chanwoo Park , Hyewon Jeong , Yik Siu Chan , Xuhai Xu , Daniel McDuff , Hyeonhoon Lee , Marzyeh Ghassemi , Cynthia Breazeal , Hae Won Park

In recent years, instruction-based image editing methods have garnered significant attention in image editing. However, despite encompassing a wide range of editing priors, these methods are helpless when handling editing tasks that are…

Graphics · Computer Science 2024-03-28 Ruoyu Zhao , Qingnan Fan , Fei Kou , Shuai Qin , Hong Gu , Wei Wu , Pengcheng Xu , Mingrui Zhu , Nannan Wang , Xinbo Gao

The recent GAN inversion methods have been able to successfully invert the real image input to the corresponding editable latent code in StyleGAN. By combining with the language-vision model (CLIP), some text-driven image manipulation…

Computer Vision and Pattern Recognition · Computer Science 2023-09-22 Yunpeng Bai , Zihan Zhong , Chao Dong , Weichen Zhang , Guowei Xu , Chun Yuan

This paper delves into the text-guided image editing task, focusing on modifying a reference image according to user-specified textual feedback to embody specific attributes. Despite recent advancements, a persistent challenge remains that…

Computer Vision and Pattern Recognition · Computer Science 2024-03-12 Lidong Zeng , Zhedong Zheng , Yinwei Wei , Tat-seng Chua

We propose a novel framework for learning high-level cognitive capabilities in robot manipulation tasks, such as making a smiley face using building blocks. These tasks often involve complex multi-step reasoning, presenting significant…

Robotics · Computer Science 2023-05-31 Chuhao Jin , Wenhui Tan , Jiange Yang , Bei Liu , Ruihua Song , Limin Wang , Jianlong Fu

Complex tasks involving tool integration pose significant challenges for Large Language Models (LLMs), leading to the emergence of multi-agent workflows as a promising solution. Reflection has emerged as an effective strategy for correcting…

Artificial Intelligence · Computer Science 2025-06-06 Zikang Guo , Benfeng Xu , Xiaorui Wang , Zhendong Mao

We introduce DriveAgent, a novel multi-agent autonomous driving framework that leverages large language model (LLM) reasoning combined with multimodal sensor fusion to enhance situational understanding and decision-making. DriveAgent…

Robotics · Computer Science 2025-05-06 Xinmeng Hou , Wuqi Wang , Long Yang , Hao Lin , Jinglun Feng , Haigen Min , Xiangmo Zhao

Engineering design problems often involve large state and action spaces along with highly sparse rewards. Since an exhaustive search of those spaces is not feasible, humans utilize relevant domain knowledge to condense the search space.…

Artificial Intelligence · Computer Science 2021-10-12 Ayush Raina , Lucas Puentes , Jonathan Cagan , Christopher McComb

Real-world multimodal applications often require any-to-any capabilities, enabling both understanding and generation across modalities including text, image, audio, and video. However, integrating the strengths of autoregressive language…

Machine Learning · Computer Science 2025-08-15 Jiulin Li , Ping Huang , Yexin Li , Shuo Chen , Juewen Hu , Ye Tian