English
Related papers

Related papers: PipeFlow: Pipelined Processing and Motion-Aware Fr…

200 papers

Recent advances in flow-based generative models have enabled training-free, text-guided image editing by inverting an image into its latent noise and regenerating it under a new target conditional guidance. However, existing methods…

Computer Vision and Pattern Recognition · Computer Science 2026-04-03 Thinh Dao , Zhen Wang , Kien T. Pham , Long Chen

Instruction-based video editing promises to democratize content creation, yet its progress is severely hampered by the scarcity of large-scale, high-quality training data. We introduce Ditto, a holistic framework designed to tackle this…

Computer Vision and Pattern Recognition · Computer Science 2025-12-18 Qingyan Bai , Qiuyu Wang , Hao Ouyang , Yue Yu , Hanlin Wang , Wen Wang , Ka Leong Cheng , Shuailei Ma , Yanhong Zeng , Zichen Liu , Yinghao Xu , Yujun Shen , Qifeng Chen

We introduce $\texttt{PairFlow}$, a lightweight preprocessing step for training Discrete Flow Models (DFMs) to achieve few-step sampling without requiring a pretrained teacher. DFMs have recently emerged as a new class of generative models…

Machine Learning · Computer Science 2026-05-26 Mingue Park , Jisung Hwang , Seungwoo Yoo , Kyeongmin Yeo , Minhyuk Sung

Modern machine learning pipelines are limited due to data availability, storage quotas, privacy regulations, and expensive annotation processes. These constraints make it difficult or impossible to train and update large-scale models on…

Computer Vision and Pattern Recognition · Computer Science 2023-04-06 Andrés Villa , Juan León Alcázar , Motasem Alfarra , Kumail Alhamoud , Julio Hurtado , Fabian Caba Heilbron , Alvaro Soto , Bernard Ghanem

Diffusion Transformers (DiTs) have gained increasing adoption in high-quality image and video generation. As demand for higher-resolution images and longer videos increases, single-GPU inference becomes inefficient due to increased latency…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-26 Jiacheng Yang , Jun Wu , Yaoyao Ding , Zhiying Xu , Yida Wang , Gennady Pekhimenko

Rectified-flow-based diffusion transformers like FLUX and OpenSora have demonstrated outstanding performance in the field of image and video generation. Despite their robust generative capabilities, these models often struggle with…

Computer Vision and Pattern Recognition · Computer Science 2025-06-16 Jiangshan Wang , Junfu Pu , Zhongang Qi , Jiayi Guo , Yue Ma , Nisha Huang , Yuxin Chen , Xiu Li , Ying Shan

Diffusion models have shown great promise for image and video generation, but sampling from state-of-the-art models requires expensive numerical integration of a generative ODE. One approach for tackling this problem is rectified flows,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Sangyun Lee , Zinan Lin , Giulia Fanti

Scene flow estimation, which aims to predict per-point 3D displacements of dynamic scenes, is a fundamental task in the computer vision field. However, previous works commonly suffer from unreliable correlation caused by locally constrained…

Computer Vision and Pattern Recognition · Computer Science 2024-05-13 Jiuming Liu , Guangming Wang , Weicai Ye , Chaokang Jiang , Jinru Han , Zhe Liu , Guofeng Zhang , Dalong Du , Hesheng Wang

Video Frame Interpolation (VFI) is a fundamental yet challenging task in computer vision, particularly under conditions involving large motion, occlusion, and lighting variation. Recent advancements in event cameras have opened up new…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Hanle Zheng , Xujie Han , Zegang Peng , Shangbin Zhang , Guangxun Du , Zhuo Zou , Xilin Wang , Jibin Wu , Hao Guo , Lei Deng

Inversion-based image editing in flow matching models has emerged as a powerful paradigm for training-free, text-guided image manipulation. A central challenge in this paradigm is the injection dilemma: injecting source features during…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Guandong Li , Zhaobin Chu

Scaling video diffusion transformers is fundamentally bottlenecked by two compounding costs: the expensive quadratic complexity of attention per step, and the iterative sampling steps. In this work, we propose EFlow, an efficient few-step…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Dogyun Park , Yanyu Li , Sergey Tulyakov , Anil Kag

Diffusion models have revolutionized the field of content synthesis and editing. Recent models have replaced the traditional UNet architecture with the Diffusion Transformer (DiT), and employed flow-matching for improved training and…

Computer Vision and Pattern Recognition · Computer Science 2025-09-10 Omri Avrahami , Or Patashnik , Ohad Fried , Egor Nemchinov , Kfir Aberman , Dani Lischinski , Daniel Cohen-Or

Long-form video processing fundamentally challenges vision-language models (VLMs) due to the high computational costs of handling extended temporal sequences. Existing token pruning and feature merging methods often sacrifice critical…

Computation and Language · Computer Science 2025-09-11 Chuanqi Cheng , Jian Guan , Wei Wu , Rui Yan

The advent of Video Diffusion Transformers (Video DiTs) marks a milestone in video generation. However, directly applying existing video editing methods to Video DiTs often incurs substantial computational overhead, due to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-30 Lingling Cai , Kang Zhao , Hangjie Yuan , Xiang Wang , Yingya Zhang , Kejie Huang

Conditional diffusion models have demonstrated impressive performance in image manipulation tasks. The general pipeline involves adding noise to the image and then denoising it. However, this method faces a trade-off problem: adding too…

Computer Vision and Pattern Recognition · Computer Science 2023-07-18 Luozhou Wang , Shuai Yang , Shu Liu , Ying-cong Chen

High-quality training data play a key role in image segmentation tasks. Usually, pixel-level annotations are expensive, laborious and time-consuming for the large volume of training data. To reduce labelling cost and improve segmentation…

Computer Vision and Pattern Recognition · Computer Science 2021-10-27 Yuying Hao , Yi Liu , Zewu Wu , Lin Han , Yizhou Chen , Guowei Chen , Lutao Chu , Shiyu Tang , Zhiliang Yu , Zeyu Chen , Baohua Lai

Large Language Models (LLMs) and Vision-Language Models (VLMs) have demonstrated remarkable reasoning and generalization capabilities in video understanding; however, their application in video editing remains largely underexplored. This…

Computer Vision and Pattern Recognition · Computer Science 2025-05-20 Yuzhi Li , Haojun Xu , Feng Tian

Recently, great progress has been achieved in text-to-video (T2V) generation by scaling transformer-based diffusion models to billions of parameters, which can generate high-quality videos. However, existing models typically produce only…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Akio Kodaira , Tingbo Hou , Ji Hou , Markos Georgopoulos , Felix Juefei-Xu , Masayoshi Tomizuka , Yue Zhao

Human motion synthesis is a fundamental task in computer animation. Recent methods based on diffusion models or GPT structure demonstrate commendable performance but exhibit drawbacks in terms of slow sampling speeds and error accumulation.…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Vincent Tao Hu , Wenzhe Yin , Pingchuan Ma , Yunlu Chen , Basura Fernando , Yuki M Asano , Efstratios Gavves , Pascal Mettes , Bjorn Ommer , Cees G. M. Snoek

Diffusion models have revolutionized text-to-image generation with its exceptional quality and creativity. However, its multi-step sampling process is known to be slow, often requiring tens of inference steps to obtain satisfactory results.…

Machine Learning · Computer Science 2024-03-26 Xingchao Liu , Xiwen Zhang , Jianzhu Ma , Jian Peng , Qiang Liu
‹ Prev 1 3 4 5 6 7 10 Next ›