English
Related papers

Related papers: GrounDiT: Grounding Diffusion Transformers via Noi…

200 papers

Controllable pathology image synthesis requires reliable regulation of spatial layout, tissue morphology, and semantic detail. However, existing text-guided diffusion models offer only coarse global control and lack the ability to enforce…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Yuntao Shou , Xiangyong Cao , Qian Zhao , Deyu Meng

Audio-driven talking head generation is critical for applications such as virtual assistants, video games, and films, where natural lip movements are essential. Despite progress in this field, challenges remain in producing both consistent…

Computer Vision and Pattern Recognition · Computer Science 2025-07-08 Yucheng Wang , Dan Xu

Diffusion models have demonstrated excellent potential for generating diverse images. However, their performance often suffers from slow generation due to iterative denoising. Knowledge distillation has been recently proposed as a remedy…

Computer Vision and Pattern Recognition · Computer Science 2023-06-12 Jiatao Gu , Shuangfei Zhai , Yizhe Zhang , Lingjie Liu , Josh Susskind

Sketch-based terrain generation seeks to create realistic landscapes for virtual environments in various applications such as computer games, animation and virtual reality. Recently, deep learning based terrain generation has emerged,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-15 Zexin Hu , Kun Hu , Clinton Mo , Lei Pan , Zhiyong Wang

Recent research arXiv:2410.15027 has explored the use of diffusion transformers (DiTs) for task-agnostic image generation by simply concatenating attention tokens across images. However, despite substantial computational resources, the…

Computer Vision and Pattern Recognition · Computer Science 2024-11-06 Lianghua Huang , Wei Wang , Zhi-Fan Wu , Yupeng Shi , Huanzhang Dou , Chen Liang , Yutong Feng , Yu Liu , Jingren Zhou

Diffusion models have demonstrated excellent capabilities in text-to-image generation. Their semantic understanding (i.e., prompt following) ability has also been greatly improved with large language models (e.g., T5, Llama). However,…

Computer Vision and Pattern Recognition · Computer Science 2024-11-05 Anthony Chen , Jianjin Xu , Wenzhao Zheng , Gaole Dai , Yida Wang , Renrui Zhang , Haofan Wang , Shanghang Zhang

Diffusion Transformers (DiTs) have greatly advanced text-to-image generation, but models still struggle to generate the correct spatial relations between objects as specified in the text prompt. In this study, we adopt a mechanistic…

Artificial Intelligence · Computer Science 2026-04-07 Binxu Wang , Jingxuan Fan , Xu Pan

The urban environment is characterized by complex spatio-temporal dynamics arising from diverse human activities and interactions. Effectively modeling these dynamics is essential for understanding and optimizing urban systems. In this…

Machine Learning · Computer Science 2025-10-21 Yuan Yuan , Chonghua Han , Jingtao Ding , Guozhen Zhang , Depeng Jin , Yong Li

Accurate channel modeling is fundamental to design and evaluation of Terahertz (THz) ultra-massive multiple-input multiple-output (UM-MIMO) systems. However, existing model-based approaches typically rely on simplified assumptions, such as…

Signal Processing · Electrical Eng. & Systems 2026-05-20 Zhengdong Hu , Chong Han

Multi-Modal Diffusion Transformers (MM-DiTs) encode rich representations for training-free concept grounding, but existing attention-based methods often produce overlapping activations on visually confusable concepts, a failure mode we call…

Computer Vision and Pattern Recognition · Computer Science 2026-05-27 Jian Zhang , Zhijun Zhang

Recent text-to-image diffusion models have demonstrated an astonishing capacity to generate high-quality images. However, researchers mainly studied the way of synthesizing images with only text prompts. While some works have explored using…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Jinheng Xie , Yuexiang Li , Yawen Huang , Haozhe Liu , Wentian Zhang , Yefeng Zheng , Mike Zheng Shou

We present VoiceDiT, a multi-modal generative model for producing environment-aware speech and audio from text and visual prompts. While aligning speech with text is crucial for intelligible speech, achieving this alignment in noisy…

Audio and Speech Processing · Electrical Eng. & Systems 2024-12-30 Jaemin Jung , Junseok Ahn , Chaeyoung Jung , Tan Dat Nguyen , Youngjoon Jang , Joon Son Chung

The diffusion model has provided a strong tool for implementing text-to-image (T2I) and image-to-image (I2I) generation. Recently, topology and texture control are popular explorations, e.g., ControlNet, IP-Adapter, Ctrl-X, and DSG. These…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Jia Li , Nan Gao , Huaibo Huang , Ran He

Diffusion models with their powerful expressivity and high sample quality have achieved State-Of-The-Art (SOTA) performance in the generative domain. The pioneering Vision Transformer (ViT) has also demonstrated strong modeling capabilities…

Computer Vision and Pattern Recognition · Computer Science 2024-08-30 Ali Hatamizadeh , Jiaming Song , Guilin Liu , Jan Kautz , Arash Vahdat

Diffusion Transformers have established a new state-of-the-art in image synthesis, but the high computational cost of iterative sampling severely hampers their practical deployment. While existing acceleration methods often focus on the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Wenhao Sun , Ji Li , Zhaoqiang Liu

Layout-to-image generation refers to the task of synthesizing photo-realistic images based on semantic layouts. In this paper, we propose LayoutDiffuse that adapts a foundational diffusion model pretrained on large-scale image or text-image…

Computer Vision and Pattern Recognition · Computer Science 2023-02-20 Jiaxin Cheng , Xiao Liang , Xingjian Shi , Tong He , Tianjun Xiao , Mu Li

In this paper, we address the task of multimodal-to-speech generation, which aims to synthesize high-quality speech from multiple input modalities: text, video, and reference audio. This task has gained increasing attention due to its wide…

Computer Vision and Pattern Recognition · Computer Science 2025-10-06 Jeongsoo Choi , Ji-Hoon Kim , Kim Sung-Bin , Tae-Hyun Oh , Joon Son Chung

In-context generation significantly enhances Diffusion Transformers (DiTs) by enabling controllable image-to-image generation through reference examples. However, the resulting input concatenation drastically increases sequence length,…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Junqing Lin , Xingyu Zheng , Pei Cheng , Bin Fu , Jingwei Sun , Guangzhong Sun

Recent advances in image generation have led to remarkable improvements in synthesizing perspective images. However, these models still struggle with panoramic image generation due to unique challenges, including varying levels of geometric…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Hakan Çapuk , Andrew Bond , Muhammed Burak Kızıl , Emir Göçen , Erkut Erdem , Aykut Erdem

Diffusion models have attracted significant attention due to the remarkable ability to create content and generate data for tasks like image classification. However, the usage of diffusion models to generate the high-quality object…

Computer Vision and Pattern Recognition · Computer Science 2024-02-20 Kai Chen , Enze Xie , Zhe Chen , Yibo Wang , Lanqing Hong , Zhenguo Li , Dit-Yan Yeung