English
Related papers

Related papers: DiTraj: training-free trajectory control for video…

200 papers

Diffusion transformer (DiT) achieves remarkable performance in visual generation, but its iterative denoising process combined with larger capacity leads to a high inference cost. Recent works have demonstrated that the iterative denoising…

Computer Vision and Pattern Recognition · Computer Science 2026-01-01 Yonglak Son , Suhyeok Kim , Seungryong Kim , Young Geun Kim

Recent advances in text-to-image (T2I) diffusion models have enabled impressive image generation capabilities guided by text prompts. However, extending these techniques to video generation remains challenging, with existing text-to-video…

Computer Vision and Pattern Recognition · Computer Science 2024-08-13 Weifeng Chen , Yatai Ji , Jie Wu , Hefeng Wu , Pan Xie , Jiashi Li , Xin Xia , Xuefeng Xiao , Liang Lin

In e-commerce and digital marketing, generating high-fidelity human-product demonstration videos is important for effective product presentation. However, most existing frameworks either fail to preserve the identities of both humans and…

Computer Vision and Pattern Recognition · Computer Science 2025-08-28 Lizhen Wang , Zhurong Xia , Tianshu Hu , Pengrui Wang , Pengfei Wei , Zerong Zheng , Ming Zhou , Yuan Zhang , Mingyuan Gao

Diffusion models exhibit excellent sample quality, but existing guidance methods often require additional model training or are limited to specific tasks. We revisit guidance in diffusion models from the perspective of variational inference…

Machine Learning · Computer Science 2025-05-27 Kushagra Pandey , Farrin Marouf Sofian , Felix Draxler , Theofanis Karaletsos , Stephan Mandt

Video try-on stands as a promising area for its tremendous real-world potential. Prior works are limited to transferring product clothing images onto person videos with simple poses and backgrounds, while underperforming on casually…

Computer Vision and Pattern Recognition · Computer Science 2024-06-10 Jun Zheng , Fuwei Zhao , Youjiang Xu , Xin Dong , Xiaodan Liang

While Diffusion Transformers (DiT) have advanced non-autoregressive (NAR) speech synthesis, their high computational demands remain an limitation. Existing DiT-based text-to-speech (TTS) model acceleration approaches mainly focus on…

Sound · Computer Science 2025-09-15 Yanru Huo , Ziyue Jiang , Zuoli Tang , Qingyang Hong , Zhou Zhao

Diffusion models show promising generation capability for a variety of data. Despite their high generation quality, the inference for diffusion models is still time-consuming due to the numerous sampling iterations required. To accelerate…

Computer Vision and Pattern Recognition · Computer Science 2023-10-26 Kexun Zhang , Xianjun Yang , William Yang Wang , Lei Li

Video diffusion models have rich world priors, but their use in spatial tasks is limited by poor control, spatial-temporal inconsistent results, and entangled scene-camera dynamics. Current approaches, such as per-task fine-tuning or…

Graphics · Computer Science 2026-03-24 Chenxi Song , Yanming Yang , Tong Zhao , Ruibo Li , Chi Zhang

The intuitive nature of drag-based interaction has led to its growing adoption for controlling object trajectories in image-to-video synthesis. Still, existing methods that perform dragging in the 2D space usually face ambiguity when…

Computer Vision and Pattern Recognition · Computer Science 2025-03-31 Hanlin Wang , Hao Ouyang , Qiuyu Wang , Wen Wang , Ka Leong Cheng , Qifeng Chen , Yujun Shen , Limin Wang

Text-driven video editing aims to modify video content based on natural language instructions. While recent training-free methods have leveraged pretrained diffusion models, they often rely on an inversion-editing paradigm. This paradigm…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Guangzhao Li , Yanming Yang , Chenxi Song , Chi Zhang

We propose TraceRL, a trajectory-aware reinforcement learning framework for diffusion language models (DLMs) that incorporates preferred inference trajectory into post-training, and is applicable across different architectures. Equipped…

Computation and Language · Computer Science 2025-09-09 Yinjie Wang , Ling Yang , Bowen Li , Ye Tian , Ke Shen , Mengdi Wang

DiT-based video generation has achieved remarkable results, but research into enhancing existing models remains relatively unexplored. In this work, we introduce a training-free approach to enhance the coherence and quality of DiT-based…

Computer Vision and Pattern Recognition · Computer Science 2025-02-28 Yang Luo , Xuanlei Zhao , Mengzhao Chen , Kaipeng Zhang , Wenqi Shao , Kai Wang , Zhangyang Wang , Yang You

Recent works such as REPA have shown that guiding diffusion models with external semantic features (e.g., DINO) can significantly accelerate the training of diffusion transformers (DiTs). However, the use of pretrained external features as…

Computer Vision and Pattern Recognition · Computer Science 2026-03-17 Lingchen Sun , Rongyuan Wu , Zhengqiang Zhang , Ruibin Li , Yujing Sun , Shuaizheng Liu , Lei Zhang

Diffusion Transformers (DiTs) have recently driven significant progress in text-to-video (T2V) generation. However, generating multiple videos with consistent characters and backgrounds remains a significant challenge. Existing methods…

Computer Vision and Pattern Recognition · Computer Science 2025-10-27 Han Yan , Xibin Song , Yifu Wang , Hongdong Li , Pan Ji , Chao Ma

Due to lack of fully publicly available text-to-video models, current video editing methods tend to build on pre-trained text-to-image generation models, however, they still face grand challenges in dealing with the local editing of video…

Computer Vision and Pattern Recognition · Computer Science 2024-09-06 Deyin Liu , Lin Yuanbo Wu , Xianghua Xie

In recent years, diffusion models have demonstrated remarkable potential across diverse domains, from vision generation to language modeling. Transferring its generative capabilities to modern end-to-end autonomous driving systems has also…

Robotics · Computer Science 2025-09-17 Xuefeng Jiang , Yuan Ma , Pengxiang Li , Leimeng Xu , Xin Wen , Kun Zhan , Zhongpu Xia , Peng Jia , Xianpeng Lang , Sheng Sun

Diffusion Transformers (DiTs) have demonstrated exceptional capabilities in text-to-image synthesis. However, in the domain of controllable text-to-image generation using DiTs, most existing methods still rely on the ControlNet paradigm…

Computer Vision and Pattern Recognition · Computer Science 2025-08-15 Shanyuan Liu , Jian Zhu , Junda Lu , Yue Gong , Liuzhuozheng Li , Bo Cheng , Yuhang Ma , Liebucha Wu , Xiaoyu Wu , Dawei Leng , Yuhui Yin

End-to-end autonomous driving systems directly generate driving policies from raw sensor inputs. While these systems can extract effective environmental features for planning, relying on auxiliary perception tasks, developing perception…

Computer Vision and Pattern Recognition · Computer Science 2025-12-02 Xingtai Gui , Jianbo Zhao , Wencheng Han , Jikai Wang , Jiahao Gong , Feiyang Tan , Cheng-zhong Xu , Jianbing Shen

We propose a diffusion-based approach for Text-to-Image (T2I) generation with interactive 3D layout control. Layout control has been widely studied to alleviate the shortcomings of T2I diffusion models in understanding objects' placement…

Computer Vision and Pattern Recognition · Computer Science 2024-08-28 Abdelrahman Eldesokey , Peter Wonka

Recent advances have reformulated diffusion models as deterministic ordinary differential equations (ODEs) through the framework of flow matching, providing a unified formulation for the noise-to-data generative process. Various…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Yan Luo , Henry Huang , Todd Y. Zhou , Mengyu Wang