中文
相关论文

相关论文: Reasoning to Align: Implicit Reasoning in Diffusio…

200 篇论文

Recently large-scale language-image models (e.g., text-guided diffusion models) have considerably improved the image generation capabilities to generate photorealistic images in various domains. Based on this success, current image editing…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Wenkai Dong , Song Xue , Xiaoyue Duan , Shumin Han

GAN inversion is indispensable for applying the powerful editability of GAN to real images. However, existing methods invert video frames individually often leading to undesired inconsistent results over time. In this paper, we propose a…

计算机视觉与模式识别 · 计算机科学 2023-08-16 Yangyang Xu , Shengfeng He , Kwan-Yee K. Wong , Ping Luo

A significant research effort is focused on exploiting the amazing capacities of pretrained diffusion models for the editing of images.They either finetune the model, or invert the image in the latent space of the pretrained model. However,…

计算机视觉与模式识别 · 计算机科学 2024-12-09 Senmao Li , Joost van de Weijer , Taihang Hu , Fahad Shahbaz Khan , Qibin Hou , Yaxing Wang , Jian Yang , Ming-Ming Cheng

Recent advances in diffusion models have successfully enabled text-guided image inpainting. While it seems straightforward to extend such editing capability into the video domain, there have been fewer works regarding text-guided video…

计算机视觉与模式识别 · 计算机科学 2024-04-02 Zhixing Zhang , Bichen Wu , Xiaoyan Wang , Yaqiao Luo , Luxin Zhang , Yinan Zhao , Peter Vajda , Dimitris Metaxas , Licheng Yu

Controllability is a fundamental requirement in video synthesis, where accurate alignment with conditioning signals is essential. Existing classifier-free guidance methods typically achieve conditioning indirectly by modeling the joint…

计算机视觉与模式识别 · 计算机科学 2025-12-25 Weiqi Li , Zehao Zhang , Liang Lin , Guangrun Wang

Visual prompt, a pair of before-and-after edited images, can convey indescribable imagery transformations and prosper in image editing. However, current visual prompt methods rely on a pretrained text-guided image-to-image generative model…

计算机视觉与模式识别 · 计算机科学 2025-01-28 Pengcheng Xu , Qingnan Fan , Fei Kou , Shuai Qin , Hong Gu , Ruoyu Zhao , Charles Ling , Boyu Wang

Diffusion Transformer(DiT) based video generation models have recently achieved impressive visual quality and temporal coherence, but they still frequently violate basic physical laws and commonsense dynamics, revealing a lack of explicit…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Selena Song , Ziming Xu , Zijun Zhang , Kun Zhou , Jiaxian Guo , Lianhui Qin , Biwei Huang

Recent advances in diffusion transformers have shown remarkable generalization in visual synthesis, yet most dense perception methods still rely on text-to-image (T2I) generators designed for stochastic generation. We revisit this paradigm…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yiqing Shi , Yiren Song , Mike Zheng Shou

While Diffusion Large Language Models (DLLMs) have demonstrated remarkable capabilities in multi-modal generation, performing precise, training-free image editing remains an open challenge. Unlike continuous diffusion models, the discrete…

计算机视觉与模式识别 · 计算机科学 2026-03-24 Zifeng Zhu , Jiaming Han , Jiaxiang Zhao , Minnan Luo , Xiangyu Yue

Diffusion Transformers (DiTs) have emerged as a leading architecture for text-to-image synthesis, producing high-quality and photorealistic images. However, the quadratic scaling properties of the attention in DiTs hinder image generation…

计算机视觉与模式识别 · 计算机科学 2025-08-12 Philipp Becker , Abhinav Mehrotra , Ruchika Chavhan , Malcolm Chadwick , Luca Morreale , Mehdi Noroozi , Alberto Gil Ramos , Sourav Bhattacharya

Most text-to-video(T2V) diffusion models depend on pre-trained text encoders for semantic alignment, yet they often fail to maintain video quality when provided with concise prompts rather than well-designed ones. The primary issue lies in…

计算机视觉与模式识别 · 计算机科学 2025-11-07 Xiangjun Zhang , Litong Gong , Yinglin Zheng , Yansong Liu , Wentao Jiang , Mingyi Xu , Biao Wang , Tiezheng Ge , Ming Zeng

Despite the significant progress in controllable music generation and editing, challenges remain in the quality and length of generated music due to the use of Mel-spectrogram representations and UNet-based model structures. To address…

音频与语音处理 · 电气工程与系统科学 2025-01-17 Siyuan Hou , Shansong Liu , Ruibin Yuan , Wei Xue , Ying Shan , Mangsuo Zhao , Chao Zhang

Recent advances in text-to-video diffusion models have enabled high-quality video synthesis, but controllable generation remains challenging, particularly under limited data and compute. Existing fine-tuning methods for conditional…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Kinam Kim , Junha Hyung , Jaegul Choo

Transformer-based diffusion models have achieved significant advancements across a variety of generative tasks. However, producing high-quality outputs typically necessitates large transformer models, which result in substantial training…

计算机视觉与模式识别 · 计算机科学 2024-12-10 Gongfan Fang , Xinyin Ma , Xinchao Wang

Recent breakthroughs in Diffusion Transformers (DiTs) have revolutionized the field of visual synthesis due to their superior scalability. To facilitate DiTs' capability of capturing meaningful internal representations, recent works such as…

计算机视觉与模式识别 · 计算机科学 2026-03-05 Mengping Yang , Zhiyu Tan , Binglei Li , Xiaomeng Yang , Hesen Chen , Hao Li

Self-reflection mechanisms that rely on purely text-based rethinking processes perform well in most multimodal tasks. However, when directly applied to long-form video understanding scenarios, they exhibit clear limitations. The fundamental…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Jiaze Li , Hao Yin , Wenhui Tan , Jingyang Chen , Boshen Xu , Yuxun Qu , Yijing Chen , Jianzhong Ju , Zhenbo Luo , Jian Luan

Music editing primarily entails the modification of instrument tracks or remixing in the whole, which offers a novel reinterpretation of the original piece through a series of operations. These music processing methods hold immense…

声音 · 计算机科学 2023-12-13 Bing Han , Junyu Dai , Weituo Hao , Xinyan He , Dong Guo , Jitong Chen , Yuxuan Wang , Yanmin Qian , Xuchen Song

Video object removal and inpainting are critical tasks in the fields of computer vision and multimedia processing, aimed at restoring missing or corrupted regions in video sequences. Traditional methods predominantly rely on flow-based…

计算机视觉与模式识别 · 计算机科学 2025-08-19 Jie Liu , Zheng Hui

Diffusion Transformers (DiT) trained with flow matching in a VAE latent space have unified visual generation across images and videos. A natural next step toward a single architecture for both generation (visual synthesis) and understanding…

Although natural language instructions offer an intuitive way to guide automated image editing, deep-learning models often struggle to achieve high-quality results, largely due to the difficulty of creating large, high-quality training…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Sherry X. Chen , Misha Sra , Pradeep Sen