English
Related papers

Related papers: Diffusion Transformers as Open-World Spatiotempora…

200 papers

The field of urban spatial-temporal prediction is advancing rapidly with the development of deep learning techniques and the availability of large-scale datasets. However, challenges persist in accessing and utilizing diverse urban…

Machine Learning · Computer Science 2024-03-08 Jiawei Jiang , Chengkai Han , Wayne Xin Zhao , Jingyuan Wang

Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to their high computational and memory costs. In this work, we present an efficient DiT…

Generating unbounded 3D scenes is crucial for large-scale scene understanding and simulation. Urban scenes, unlike natural landscapes, consist of various complex man-made objects and structures such as roads, traffic signs, vehicles, and…

Computer Vision and Pattern Recognition · Computer Science 2024-03-20 Junge Zhang , Qihang Zhang , Li Zhang , Ramana Rao Kompella , Gaowen Liu , Bolei Zhou

As deep learning technology advances and more urban spatial-temporal data accumulates, an increasing number of deep learning models are being proposed to solve urban spatial-temporal prediction problems. However, there are limitations in…

Machine Learning · Computer Science 2024-03-08 Jiawei Jiang , Chengkai Han , Wenjun Jiang , Wayne Xin Zhao , Jingyuan Wang

Removing various degradations from damaged documents greatly benefits digitization, downstream document analysis, and readability. Previous methods often treat each restoration task independently with dedicated models, leading to a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-07 Fangmin Zhao , Weichao Zeng , Zhenhang Li , Dongbao Yang , Binbin Li , Xiaojun Bi , Yu Zhou

Diffusion models have demonstrated highly-expressive generative capabilities in vision and NLP. Recent studies in reinforcement learning (RL) have shown that diffusion models are also powerful in modeling complex policies or trajectories in…

Machine Learning · Computer Science 2023-10-11 Haoran He , Chenjia Bai , Kang Xu , Zhuoran Yang , Weinan Zhang , Dong Wang , Bin Zhao , Xuelong Li

In this paper, we introduce a novel approach to trajectory generation for autonomous driving, combining the strengths of Diffusion models and Transformers. First, we use the historical trajectory data for efficient preprocessing and…

Robotics · Computer Science 2024-05-07 Chen Yang , Tianyu Shi

Diffusion Transformers (DiT) have emerged as a widely adopted backbone for high-fidelity image and video generation, yet their iterative denoising process incurs high computational costs. Existing training-free acceleration methods rely on…

Computer Vision and Pattern Recognition · Computer Science 2026-02-23 Hanshuai Cui , Zhiqing Tang , Qianli Ma , Zhi Yao , Weijia Jia

In this work, we empirically study Diffusion Transformers (DiTs) for text-to-image generation, focusing on architectural choices, text-conditioning strategies, and training protocols. We evaluate a range of DiT-based…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Chen Chen , Rui Qian , Wenze Hu , Tsu-Jui Fu , Jialing Tong , Xinze Wang , Lezhi Li , Bowen Zhang , Alex Schwing , Wei Liu , Yinfei Yang

Diffusion Transformers (DiTs) have recently achieved remarkable success in text-guided image generation. In image editing, DiTs project text and image inputs to a joint latent space, from which they decode and synthesize new images.…

Computer Vision and Pattern Recognition · Computer Science 2024-11-14 Zitao Shuai , Chenwei Wu , Zhengxu Tang , Bowen Song , Liyue Shen

Modeling and reconstructing multidimensional physical dynamics from sparse and off-grid observations presents a fundamental challenge in scientific research. Recently, diffusion-based generative modeling shows promising potential for…

Machine Learning · Computer Science 2025-09-30 Panqi Chen , Yifan Sun , Lei Cheng , Yang Yang , Weichang Li , Yang Liu , Weiqing Liu , Jiang Bian , Shikai Fang

We present Prompt Diffusion, a framework for enabling in-context learning in diffusion-based generative models. Given a pair of task-specific example images, such as depth from/to image and scribble from/to image, and a text guidance, our…

Computer Vision and Pattern Recognition · Computer Science 2023-10-20 Zhendong Wang , Yifan Jiang , Yadong Lu , Yelong Shen , Pengcheng He , Weizhu Chen , Zhangyang Wang , Mingyuan Zhou

Diffusion Transformers have established a new state-of-the-art in image synthesis, but the high computational cost of iterative sampling severely hampers their practical deployment. While existing acceleration methods often focus on the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Wenhao Sun , Ji Li , Zhaoqiang Liu

One of the primary objectives of satellite remote sensing is to capture the complex dynamics of the Earth environment, which encompasses tasks such as reconstructing continuous cloud-free image sequences, detecting land cover changes, and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Yuxiang Zhang , Shunlin Liang , Wenyuan Li , Han Ma , Jianglei Xu , Yichuan Ma , Jiangwei Xie , Wei Li , Mengmeng Zhang , Ran Tao , Xiang-Gen Xia

Recent advancements in video diffusion models based on Diffusion Transformers (DiTs) have achieved remarkable success in generating temporally coherent videos. Yet, a fundamental question persists: how do these models internally establish…

Computer Vision and Pattern Recognition · Computer Science 2025-06-24 Jisu Nam , Soowon Son , Dahyun Chung , Jiyoung Kim , Siyoon Jin , Junhwa Hur , Seungryong Kim

Efficiently modeling spatio-temporal (ST) physical processes and observations presents a challenging problem for the deep learning community. Many recent studies have concentrated on meticulously reconciling various advantages, leading to…

Artificial Intelligence · Computer Science 2024-06-04 Hao Wu , Yuxuan Liang , Wei Xiong , Zhengyang Zhou , Wei Huang , Shilong Wang , Kun Wang

Recently, the advent of generative AI technologies has made transformational impacts on our daily lives, yet its application in scientific applications remains in its early stages. Data scarcity is a major, well-known barrier in data-driven…

Machine Learning · Computer Science 2025-01-03 Junhuan Yang , Yuzhou Zhang , Yi Sheng , Youzuo Lin , Lei Yang

This study introduces a novel Remote Sensing (RS) Urban Prediction (UP) task focused on future urban planning, which aims to forecast urban layouts by utilizing information from existing urban layouts and planned change maps. To address the…

Computer Vision and Pattern Recognition · Computer Science 2024-07-18 Zeyu Wang , Zecheng Hao , Jingyu Lin , Yuchao Feng , Yufei Guo

Content-aware layout generation is a critical task in graphic design automation, focused on creating visually appealing arrangements of elements that seamlessly blend with a given background image. The variety of real-world applications…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Zeyang Liu , Le Wang , Sanping Zhou , Yuxuan Wu , Xiaolong Sun , Gang Hua , Haoxiang Li

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge maps. This multimodal…

Computer Vision and Pattern Recognition · Computer Science 2026-04-01 Bharath Krishnamurthy , Ajita Rattani