中文
相关论文

相关论文: UniMC: Taming Diffusion Transformer for Unified Ke…

200 篇论文

We present UniFluid, a unified autoregressive framework for joint visual generation and understanding leveraging continuous visual tokens. Our unified autoregressive architecture processes multimodal image and text inputs, generating…

The task of layout-to-image generation involves synthesizing images based on the captions of objects and their spatial positions. Existing methods still struggle in complex layout generation, where common bad cases include object missing,…

计算机视觉与模式识别 · 计算机科学 2024-10-21 Bo Cheng , Yuhang Ma , Liebucha Wu , Shanyuan Liu , Ao Ma , Xiaoyu Wu , Dawei Leng , Yuhui Yin

We present ILLUME+ that leverages dual visual tokenization and a diffusion decoder to improve both deep semantic understanding and high-fidelity image generation. Existing unified models have struggled to simultaneously handle the three…

计算机视觉与模式识别 · 计算机科学 2025-04-04 Runhui Huang , Chunwei Wang , Junwei Yang , Guansong Lu , Yunlong Yuan , Jianhua Han , Lu Hou , Wei Zhang , Lanqing Hong , Hengshuang Zhao , Hang Xu

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with…

计算机视觉与模式识别 · 计算机科学 2025-10-01 Yuxin Song , Wenkai Dong , Shizun Wang , Qi Zhang , Song Xue , Tao Yuan , Hu Yang , Haocheng Feng , Hang Zhou , Xinyan Xiao , Jingdong Wang

Recent multimodal image generators such as GPT-4o, Gemini 2.0 Flash, and Gemini 2.5 Pro excel at following complex instructions, editing images and maintaining concept consistency. However, they are still evaluated by disjoint toolkits:…

计算机视觉与模式识别 · 计算机科学 2025-05-29 Hang Hua , Ziyun Zeng , Yizhi Song , Yunlong Tang , Liu He , Daniel Aliaga , Wei Xiong , Jiebo Luo

Although recent advances in visual generation have been remarkable, most existing architectures still depend on distinct encoders for images and text. This separation constrains diffusion models' ability to perform cross-modal reasoning and…

计算机视觉与模式识别 · 计算机科学 2025-10-15 Kevin Li , Manuel Brack , Sudeep Katakol , Hareesh Ravi , Ajinkya Kale

Existing diffusion-based 3D scene generation methods primarily operate in 2D image/video latent spaces, which makes maintaining cross-view appearance and geometric consistency inherently challenging. To bridge this gap, we present OneWorld,…

计算机视觉与模式识别 · 计算机科学 2026-03-18 Sensen Gao , Zhaoqing Wang , Qihang Cao , Dongdong Yu , Changhu Wang , Tongliang Liu , Mingming Gong , Jiawang Bian

Diffusion probabilistic models (DPMs) have demonstrated a very promising ability in high-resolution image synthesis. However, sampling from a pre-trained DPM is time-consuming due to the multiple evaluations of the denoising network, making…

机器学习 · 计算机科学 2023-10-18 Wenliang Zhao , Lujia Bai , Yongming Rao , Jie Zhou , Jiwen Lu

The image compression model has long struggled with adaptability and generalization, as the decoded bitstream typically serves only human or machine needs and fails to preserve information for unseen visual tasks. Therefore, this paper…

计算机视觉与模式识别 · 计算机科学 2025-01-09 Kangsheng Yin , Quan Liu , Xuelin Shen , Yulin He , Wenhan Yang , Shiqi Wang

Recent advances in text-to-image diffusion models have enabled the generation of diverse and high-quality images. While impressive, the images often fall short of depicting subtle details and are susceptible to errors due to ambiguity in…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Idan Schwartz , Vésteinn Snæbjarnarson , Hila Chefer , Ryan Cotterell , Serge Belongie , Lior Wolf , Sagie Benaim

Camouflage Images Generation (CIG) is an emerging research area that focuses on synthesizing images in which objects are harmoniously blended and exhibit high visual consistency with their surroundings. Existing methods perform CIG by…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Yuhang Qian , Haiyan Chen , Wentong Li , Ningzhong Liu , Jie Qin

Abstract. The advancement of deep learning has coincided with the proliferation of both models and available data. The surge in dataset sizes and the subsequent surge in computational requirements have led to the development of the Dataset…

计算机视觉与模式识别 · 计算机科学 2024-09-24 Jun-Yeong Moon , Jung Uk Kim , Gyeong-Moon Park

Multi-modal object tracking integrates auxiliary modalities such as depth, thermal infrared, event flow, and language to provide additional information beyond RGB images, showing great potential in improving tracking stabilization in…

计算机视觉与模式识别 · 计算机科学 2025-05-20 Shiyu Xuan , Zechao Li , Jinhui Tang

Recent advancements in image generative foundation models have prioritized quality improvements but often at the cost of increased computational complexity and inference latency. To address this critical trade-off, we introduce HiDream-I1,…

The performance of unified multimodal models for image generation and editing is fundamentally constrained by the quality and comprehensiveness of their training data. While existing datasets have covered basic tasks like style transfer and…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Zhihong Chen , Xuehai Bai , Yang Shi , Chaoyou Fu , Huanyu Zhang , Haotian Wang , Xiaoyan Sun , Zhang Zhang , Liang Wang , Yuanxing Zhang , Pengfei Wan , Yi-Fan Zhang

Generative models trained on internet-scale data are capable of generating novel and realistic texts, images, and videos. A natural next question is whether these models can advance science, for example by generating novel stable materials.…

机器学习 · 计算机科学 2024-06-05 Sherry Yang , KwangHwan Cho , Amil Merchant , Pieter Abbeel , Dale Schuurmans , Igor Mordatch , Ekin Dogus Cubuk

Motivated by discrete diffusion's success in language-vision modeling, we explore its potential for multi-view generation, a task dominated by continuous approaches. We introduce ViewMask-1-to-3, formulating multi-view synthesis as a…

计算机视觉与模式识别 · 计算机科学 2026-03-16 Ruishu Zhu , Zhihao Huang , Jiacheng Sun , Ping Luo , Hongyuan Zhang , Xuelong Li

Current unified multimodal models for image generation and editing typically rely on massive parameter scales (e.g., >10B), entailing prohibitive training costs and deployment footprints. In this work, we present DeepGen 1.0, a lightweight…

Multi-modal magnetic resonance imaging (MRI) provides rich, complementary information for analyzing diseases. However, the practical challenges of acquiring multiple MRI modalities, such as cost, scan time, and safety considerations, often…

图像与视频处理 · 电气工程与系统科学 2024-09-16 Zhaohu Xing , Sicheng Yang , Sixiang Chen , Tian Ye , Yijun Yang , Jing Qin , Lei Zhu

Modern image retrieval methods typically rely on fine-tuning pre-trained encoders to extract image-level descriptors. However, the most widely used models are pre-trained on ImageNet-1K with limited classes. The pre-trained feature…

计算机视觉与模式识别 · 计算机科学 2023-04-13 Xiang An , Jiankang Deng , Kaicheng Yang , Jaiwei Li , Ziyong Feng , Jia Guo , Jing Yang , Tongliang Liu