English
Related papers

Related papers: PipeDiT: Accelerating Diffusion Transformers in Vi…

200 papers

Diffusion models have achieved great success in synthesizing high-quality images. However, generating high-resolution images with diffusion models is still challenging due to the enormous computational costs, resulting in a prohibitive…

Computer Vision and Pattern Recognition · Computer Science 2024-07-16 Muyang Li , Tianle Cai , Jiaxin Cao , Qinsheng Zhang , Han Cai , Junjie Bai , Yangqing Jia , Ming-Yu Liu , Kai Li , Song Han

Generating multiview images from a single view facilitates the rapid generation of a 3D mesh conditioned on a single image. Recent methods that introduce 3D global representation into diffusion models have shown the potential to generate…

Computer Vision and Pattern Recognition · Computer Science 2024-04-03 Zehuan Huang , Hao Wen , Junting Dong , Yaohui Wang , Yangguang Li , Xinyuan Chen , Yan-Pei Cao , Ding Liang , Yu Qiao , Bo Dai , Lu Sheng

Existing text-video retrieval solutions are, in essence, discriminant models focused on maximizing the conditional likelihood, i.e., p(candidates|query). While straightforward, this de facto paradigm overlooks the underlying data…

Computer Vision and Pattern Recognition · Computer Science 2023-08-22 Peng Jin , Hao Li , Zesen Cheng , Kehan Li , Xiangyang Ji , Chang Liu , Li Yuan , Jie Chen

DiT models have achieved great success in text-to-video generation, leveraging their scalability in model capacity and data scale. High content and motion fidelity aligned with text prompts, however, often require large model parameters and…

Computer Vision and Pattern Recognition · Computer Science 2026-02-02 Shilong Zhang , Wenbo Li , Shoufa Chen , Chongjian GE , Peize Sun , Yifu Zhang , Yi Jiang , Zehuan Yuan , Bingyue Peng , Ping Luo

The high memory and computation demand of large language models (LLMs) makes them challenging to be deployed on consumer devices due to limited GPU memory. Offloading can mitigate the memory constraint but often suffers from low GPU…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-06-16 Yangyijian Liu , Jun Li , Wu-Jun Li

Recent advancements in diffusion techniques have propelled image and video generation to unprecedented levels of quality, significantly accelerating the deployment and application of generative AI. However, 3D shape generation technology…

Computer Vision and Pattern Recognition · Computer Science 2025-03-28 Yangguang Li , Zi-Xin Zou , Zexiang Liu , Dehu Wang , Yuan Liang , Zhipeng Yu , Xingchao Liu , Yuan-Chen Guo , Ding Liang , Wanli Ouyang , Yan-Pei Cao

Diffusion models represent a powerful family of generative models widely used for image and video generation. However, the time-consuming deployment, long inference time, and requirements on large memory hinder their applications on…

Machine Learning · Computer Science 2025-04-18 Kafeng Wang , Jianfei Chen , He Li , Zhenpeng Mi , Jun Zhu

We present VIDIM, a generative model for video interpolation, which creates short videos given a start and end frame. In order to achieve high fidelity and generate motions unseen in the input data, VIDIM uses cascaded diffusion models to…

Computer Vision and Pattern Recognition · Computer Science 2024-04-02 Siddhant Jain , Daniel Watson , Eric Tabellion , Aleksander Hołyński , Ben Poole , Janne Kontkanen

Video diffusion generation suffers from critical sampling efficiency bottlenecks, particularly for large-scale models and long contexts. Existing video acceleration methods, adapted from image-based techniques, lack a single-step…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Jiaxiang Cheng , Bing Ma , Xuhua Ren , Hongyi Henry Jin , Kai Yu , Peng Zhang , Wenyue Li , Yuan Zhou , Tianxiang Zheng , Qinglin Lu

In recent years, there has been a significant surge of interest in unifying image comprehension and generation within Large Language Models (LLMs). This growing interest has prompted us to explore extending this unification to videos. The…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Yuying Ge , Yizhuo Li , Yixiao Ge , Ying Shan

Diffusion pipelines, renowned for their powerful visual generation capabilities, have seen widespread adoption in generative vision tasks (e.g., text-to-image/video). These pipelines typically follow an encode--diffuse--decode three-stage…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-06 Yifei Xia , Fangcheng Fu , Hao Yuan , Hanke Zhang , Xupeng Miao , Yijun Liu , Suhan Ling , Jie Jiang , Bin Cui

Diffusion Transformers (DiTs) dominate video generation but their high computational cost severely limits real-world applicability, usually requiring tens of minutes to generate a few seconds of video even on high-performance GPUs. This…

Computer Vision and Pattern Recognition · Computer Science 2025-04-29 Haocheng Xi , Shuo Yang , Yilong Zhao , Chenfeng Xu , Muyang Li , Xiuyu Li , Yujun Lin , Han Cai , Jintao Zhang , Dacheng Li , Jianfei Chen , Ion Stoica , Kurt Keutzer , Song Han

We present a new, fast and flexible pipeline for indoor scene synthesis that is based on deep convolutional generative models. Our method operates on a top-down image-based representation, and inserts objects iteratively into the scene by…

Computer Vision and Pattern Recognition · Computer Science 2018-12-03 Daniel Ritchie , Kai Wang , Yu-an Lin

This paper presents PipeBoost, a low-latency LLM serving system for multi-GPU (serverless) clusters, which can rapidly launch inference services in response to bursty requests without preemptively over-provisioning GPUs. Many LLM inference…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-03-25 Chongpeng Liu , Xiaojian Liao , Hancheng Liu , Limin Xiao , Jianxin Li

Diffusion-based generative models have demonstrated exceptional promise in the video super-resolution (VSR) task, achieving a substantial advancement in detail generation relative to prior methods. However, these approaches face significant…

Computer Vision and Pattern Recognition · Computer Science 2025-07-01 Zhongdao Wang , Guodongfang Zhao , Jingjing Ren , Bailan Feng , Shifeng Zhang , Wenbo Li

The diffusion model performs remarkable in generating high-dimensional content but is computationally intensive, especially during training. We propose Progressive Growing of Diffusion Autoencoder (PaGoDA), a novel pipeline that reduces the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Dongjun Kim , Chieh-Hsin Lai , Wei-Hsiang Liao , Yuhta Takida , Naoki Murata , Toshimitsu Uesaka , Yuki Mitsufuji , Stefano Ermon

Diffusion Transformer has shown remarkable abilities in generating high-fidelity videos, delivering visually coherent frames and rich details over extended durations. However, existing video generation models still fall short in…

Computer Vision and Pattern Recognition · Computer Science 2026-03-04 Zhaoyang Li , Dongjun Qian , Kai Su , Qishuai Diao , Xiangyang Xia , Chang Liu , Wenfei Yang , Tianzhu Zhang , Zehuan Yuan

In this paper, we introduce a novel approach to trajectory generation for autonomous driving, combining the strengths of Diffusion models and Transformers. First, we use the historical trajectory data for efficient preprocessing and…

Robotics · Computer Science 2024-05-07 Chen Yang , Tianyu Shi

Deep learning video analytic systems process live video feeds from multiple cameras with computer vision models deployed on edge or cloud. To optimize utility for these systems, which usually corresponds to query accuracy, efficient…

Networking and Internet Architecture · Computer Science 2023-06-28 Hongpeng Guo , Beitong Tian , Zhe Yang , Bo Chen , Qian Zhou , Shengzhong Liu , Klara Nahrstedt , Claudiu Danilov

Human motion video generation has advanced significantly, while existing methods still struggle with accurately rendering detailed body parts like hands and faces, especially in long sequences and intricate motions. Current approaches also…

Computer Vision and Pattern Recognition · Computer Science 2025-08-05 Qijun Gan , Yi Ren , Chen Zhang , Zhenhui Ye , Pan Xie , Xiang Yin , Zehuan Yuan , Bingyue Peng , Jianke Zhu