中文
相关论文

相关论文: Diffusion Transformers as Open-World Spatiotempora…

200 篇论文

The field of urban spatial-temporal prediction is advancing rapidly with the development of deep learning techniques and the availability of large-scale datasets. However, challenges persist in accessing and utilizing diverse urban…

机器学习 · 计算机科学 2024-03-08 Jiawei Jiang , Chengkai Han , Wayne Xin Zhao , Jingyuan Wang

Recent advances in diffusion transformers (DiTs) have set new standards in image generation, yet remain impractical for on-device deployment due to their high computational and memory costs. In this work, we present an efficient DiT…

Generating unbounded 3D scenes is crucial for large-scale scene understanding and simulation. Urban scenes, unlike natural landscapes, consist of various complex man-made objects and structures such as roads, traffic signs, vehicles, and…

计算机视觉与模式识别 · 计算机科学 2024-03-20 Junge Zhang , Qihang Zhang , Li Zhang , Ramana Rao Kompella , Gaowen Liu , Bolei Zhou

As deep learning technology advances and more urban spatial-temporal data accumulates, an increasing number of deep learning models are being proposed to solve urban spatial-temporal prediction problems. However, there are limitations in…

机器学习 · 计算机科学 2024-03-08 Jiawei Jiang , Chengkai Han , Wenjun Jiang , Wayne Xin Zhao , Jingyuan Wang

Removing various degradations from damaged documents greatly benefits digitization, downstream document analysis, and readability. Previous methods often treat each restoration task independently with dedicated models, leading to a…

计算机视觉与模式识别 · 计算机科学 2025-08-07 Fangmin Zhao , Weichao Zeng , Zhenhang Li , Dongbao Yang , Binbin Li , Xiaojun Bi , Yu Zhou

Diffusion models have demonstrated highly-expressive generative capabilities in vision and NLP. Recent studies in reinforcement learning (RL) have shown that diffusion models are also powerful in modeling complex policies or trajectories in…

机器学习 · 计算机科学 2023-10-11 Haoran He , Chenjia Bai , Kang Xu , Zhuoran Yang , Weinan Zhang , Dong Wang , Bin Zhao , Xuelong Li

In this paper, we introduce a novel approach to trajectory generation for autonomous driving, combining the strengths of Diffusion models and Transformers. First, we use the historical trajectory data for efficient preprocessing and…

机器人学 · 计算机科学 2024-05-07 Chen Yang , Tianyu Shi

Diffusion Transformers (DiT) have emerged as a widely adopted backbone for high-fidelity image and video generation, yet their iterative denoising process incurs high computational costs. Existing training-free acceleration methods rely on…

计算机视觉与模式识别 · 计算机科学 2026-02-23 Hanshuai Cui , Zhiqing Tang , Qianli Ma , Zhi Yao , Weijia Jia

In this work, we empirically study Diffusion Transformers (DiTs) for text-to-image generation, focusing on architectural choices, text-conditioning strategies, and training protocols. We evaluate a range of DiT-based…

计算机视觉与模式识别 · 计算机科学 2025-03-18 Chen Chen , Rui Qian , Wenze Hu , Tsu-Jui Fu , Jialing Tong , Xinze Wang , Lezhi Li , Bowen Zhang , Alex Schwing , Wei Liu , Yinfei Yang

Diffusion Transformers (DiTs) have recently achieved remarkable success in text-guided image generation. In image editing, DiTs project text and image inputs to a joint latent space, from which they decode and synthesize new images.…

计算机视觉与模式识别 · 计算机科学 2024-11-14 Zitao Shuai , Chenwei Wu , Zhengxu Tang , Bowen Song , Liyue Shen

Modeling and reconstructing multidimensional physical dynamics from sparse and off-grid observations presents a fundamental challenge in scientific research. Recently, diffusion-based generative modeling shows promising potential for…

机器学习 · 计算机科学 2025-09-30 Panqi Chen , Yifan Sun , Lei Cheng , Yang Yang , Weichang Li , Yang Liu , Weiqing Liu , Jiang Bian , Shikai Fang

We present Prompt Diffusion, a framework for enabling in-context learning in diffusion-based generative models. Given a pair of task-specific example images, such as depth from/to image and scribble from/to image, and a text guidance, our…

计算机视觉与模式识别 · 计算机科学 2023-10-20 Zhendong Wang , Yifan Jiang , Yadong Lu , Yelong Shen , Pengcheng He , Weizhu Chen , Zhangyang Wang , Mingyuan Zhou

Diffusion Transformers have established a new state-of-the-art in image synthesis, but the high computational cost of iterative sampling severely hampers their practical deployment. While existing acceleration methods often focus on the…

计算机视觉与模式识别 · 计算机科学 2026-03-19 Wenhao Sun , Ji Li , Zhaoqiang Liu

One of the primary objectives of satellite remote sensing is to capture the complex dynamics of the Earth environment, which encompasses tasks such as reconstructing continuous cloud-free image sequences, detecting land cover changes, and…

计算机视觉与模式识别 · 计算机科学 2026-03-09 Yuxiang Zhang , Shunlin Liang , Wenyuan Li , Han Ma , Jianglei Xu , Yichuan Ma , Jiangwei Xie , Wei Li , Mengmeng Zhang , Ran Tao , Xiang-Gen Xia

Recent advancements in video diffusion models based on Diffusion Transformers (DiTs) have achieved remarkable success in generating temporally coherent videos. Yet, a fundamental question persists: how do these models internally establish…

计算机视觉与模式识别 · 计算机科学 2025-06-24 Jisu Nam , Soowon Son , Dahyun Chung , Jiyoung Kim , Siyoon Jin , Junhwa Hur , Seungryong Kim

Efficiently modeling spatio-temporal (ST) physical processes and observations presents a challenging problem for the deep learning community. Many recent studies have concentrated on meticulously reconciling various advantages, leading to…

人工智能 · 计算机科学 2024-06-04 Hao Wu , Yuxuan Liang , Wei Xiong , Zhengyang Zhou , Wei Huang , Shilong Wang , Kun Wang

Recently, the advent of generative AI technologies has made transformational impacts on our daily lives, yet its application in scientific applications remains in its early stages. Data scarcity is a major, well-known barrier in data-driven…

机器学习 · 计算机科学 2025-01-03 Junhuan Yang , Yuzhou Zhang , Yi Sheng , Youzuo Lin , Lei Yang

This study introduces a novel Remote Sensing (RS) Urban Prediction (UP) task focused on future urban planning, which aims to forecast urban layouts by utilizing information from existing urban layouts and planned change maps. To address the…

计算机视觉与模式识别 · 计算机科学 2024-07-18 Zeyu Wang , Zecheng Hao , Jingyu Lin , Yuchao Feng , Yufei Guo

Content-aware layout generation is a critical task in graphic design automation, focused on creating visually appealing arrangements of elements that seamlessly blend with a given background image. The variety of real-world applications…

计算机视觉与模式识别 · 计算机科学 2025-12-10 Zeyang Liu , Le Wang , Sanping Zhou , Yuxuan Wu , Xiaolong Sun , Gang Hua , Haoxiang Li

Recent multimodal face generation models address the spatial control limitations of text-to-image diffusion models by augmenting text-based conditioning with spatial priors such as segmentation masks, sketches, or edge maps. This multimodal…

计算机视觉与模式识别 · 计算机科学 2026-04-01 Bharath Krishnamurthy , Ajita Rattani