English
Related papers

Related papers: ComfyMind: Toward General-Purpose Generation via T…

200 papers

Generative AI (GenAI) models have become more capable than ever at augmenting productivity and cognition across diverse contexts. However, a fundamental challenge remains as users struggle to anticipate what AI will generate. As a result,…

Human-Computer Interaction · Computer Science 2025-04-17 Bryan Min , Haijun Xia

Recent neural generation systems have demonstrated the potential for procedurally generating game content, images, stories, and more. However, most neural generation algorithms are "uncontrolled" in the sense that the user has little say in…

Artificial Intelligence · Computer Science 2022-08-08 Zhiyu Lin , Rohan Agarwal , Mark Riedl

Automatic image generation is no longer just of interest to researchers, but also to practitioners. However, current models are sensitive to the settings used and automatic optimization methods often require human involvement. To bridge…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Dominik Sobania , Martin Briesch , Franz Rothlauf

This paper presents a real-time generative drawing system that interprets and integrates both formal intent - the structural, compositional, and stylistic attributes of a sketch - and contextual intent - the semantic and thematic meaning…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Jookyung Song , Mookyoung Kang , Nojun Kwak

Despite impressive progress in high-fidelity image synthesis, generative models still struggle with logic-intensive instruction following, exposing a persistent reasoning--execution gap. Meanwhile, closed-source systems (e.g., Nano Banana)…

Computer Vision and Pattern Recognition · Computer Science 2026-04-06 Sashuai Zhou , Qiang Zhou , Jijin Hu , Hanqing Yang , Yue Cao , Junpeng Ma , Yinchao Ma , Jun Song , Tiezheng Ge , Cheng Yu , Bo Zheng , Zhou Zhao

The rapid advancement of generative models has empowered modern AI systems to comprehend and produce highly sophisticated content, even achieving human-level performance in specific domains. However, these models are fundamentally…

The practical use of text-to-image generation has evolved from simple, monolithic models to complex workflows that combine multiple specialized components. While workflow-based approaches can lead to improved image quality, crafting…

Computer Vision and Pattern Recognition · Computer Science 2024-10-03 Rinon Gal , Adi Haviv , Yuval Alaluf , Amit H. Bermano , Daniel Cohen-Or , Gal Chechik

Large language models (LLMs) have demonstrated impressive capabilities in natural language understanding and generation, including multi-step reasoning such as mathematical proving. However, existing approaches often lack an explicit and…

Computation and Language · Computer Science 2026-05-19 Yutong Li , Yitian Zhou , Xudong Wang , GuoChen , Caiyan Qin

We present Composable Diffusion (CoDi), a novel generative model capable of generating any combination of output modalities, such as language, image, video, or audio, from any combination of input modalities. Unlike existing generative AI…

Computer Vision and Pattern Recognition · Computer Science 2023-05-22 Zineng Tang , Ziyi Yang , Chenguang Zhu , Michael Zeng , Mohit Bansal

The integration of AI assistants into software development workflows is rapidly evolving, shifting from automation-assisted tasks to collaborative interactions between developers and AI. Large Language Models (LLMs) have demonstrated their…

Software Engineering · Computer Science 2025-06-16 Benedetta Donato , Leonardo Mariani , Daniela Micucci , Oliviero Riganelli , Marco Somaschini

Foundation Models (FMs) have shown remarkable capabilities in various natural language tasks. However, their ability to accurately capture stakeholder requirements remains a significant challenge for using FMs for software development. This…

Software Engineering · Computer Science 2025-05-29 Keheliya Gallaba , Ali Arabat , Dayi Lin , Mohammed Sayagh , Ahmed E. Hassan

In this work, we introduce OmniGen2, a versatile and open-source generative model designed to provide a unified solution for diverse generation tasks, including text-to-image, image editing, and in-context generation. Unlike OmniGen v1,…

In software development, the raw requirements proposed by users are frequently incomplete, which impedes the complete implementation of application functionalities. With the emergence of large language models, recent methods with the…

Software Engineering · Computer Science 2024-11-11 Sai Zhang , Zhenchang Xing , Ronghui Guo , Fangzhou Xu , Lei Chen , Zhaoyuan Zhang , Xiaowang Zhang , Zhiyong Feng , Zhiqiang Zhuang

Unified large multimodal models (LMMs) have achieved remarkable progress in general-purpose multimodal understanding and generation. However, they still operate under a ``one-size-fits-all'' paradigm and struggle to model user-specific…

Computer Vision and Pattern Recognition · Computer Science 2026-01-13 Yu Zhong , Tianwei Lin , Ruike Zhu , Yuqian Yuan , Haoyu Zheng , Liang Liang , Wenqiao Zhang , Feifei Shao , Haoyuan Li , Wanggui He , Hao Jiang , Yueting Zhuang

We present Thinking with Generated Images, a novel paradigm that fundamentally transforms how large multimodal models (LMMs) engage with visual reasoning by enabling them to natively think across text and vision modalities through…

Computer Vision and Pattern Recognition · Computer Science 2025-05-29 Ethan Chern , Zhulin Hu , Steffi Chern , Siqi Kou , Jiadi Su , Yan Ma , Zhijie Deng , Pengfei Liu

The emergence of Large Language Models (LLMs) has unified language generation tasks and revolutionized human-machine interaction. However, in the realm of image generation, a unified model capable of handling various tasks within a single…

Computer Vision and Pattern Recognition · Computer Science 2024-11-22 Shitao Xiao , Yueze Wang , Junjie Zhou , Huaying Yuan , Xingrun Xing , Ruiran Yan , Chaofan Li , Shuting Wang , Tiejun Huang , Zheng Liu

Generative retrieval is an emerging approach in information retrieval that generates identifiers (IDs) of target data based on a query, providing an efficient alternative to traditional embedding-based retrieval methods. However, existing…

Information Retrieval · Computer Science 2025-06-09 Sungyeon Kim , Xinliang Zhu , Xiaofan Lin , Muhammet Bastan , Douglas Gray , Suha Kwak

In this report, we present OpenUni, a simple, lightweight, and fully open-source baseline for unifying multimodal understanding and generation. Inspired by prevailing practices in unified model learning, we adopt an efficient training…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Size Wu , Zhonghua Wu , Zerui Gong , Qingyi Tao , Sheng Jin , Qinyue Li , Wei Li , Chen Change Loy

Human generation has achieved significant progress. Nonetheless, existing methods still struggle to synthesize specific regions such as faces and hands. We argue that the main reason is rooted in the training data. A holistic human dataset…

Computer Vision and Pattern Recognition · Computer Science 2023-09-26 Jianglin Fu , Shikai Li , Yuming Jiang , Kwan-Yee Lin , Wayne Wu , Ziwei Liu

Recent diffusion models achieve strong photorealism and fluency in video generation, yet remain fragile under abstract, sparse or complex conditions, leading to poor performance in professional production workflows such as storyboard…

Computer Vision and Pattern Recognition · Computer Science 2026-05-20 Hongji Yang , Songlian Li , Yucheng Zhou , Xiaotong Zhao , Alan Zhao , Chengzhong Xu , Jianbing Shen