中文
相关论文

相关论文: Region-Constraint In-Context Generation for Instru…

200 篇论文

Recent advances in foundation models highlight a clear trend toward unification and scaling, showing emergent capabilities across diverse domains. While image generation and editing have rapidly transitioned from task-specific to unified…

计算机视觉与模式识别 · 计算机科学 2026-04-21 Xuan Ju , Tianyu Wang , Yuqian Zhou , He Zhang , Qing Liu , Nanxuan Zhao , Zhifei Zhang , Yijun Li , Yuanhao Cai , Shaoteng Liu , Daniil Pakhomov , Zhe Lin , Soo Ye Kim , Qiang Xu

Generating customized content in videos has received increasing attention recently. However, existing works primarily focus on customized text-to-video generation for single subject, suffering from subject-missing and attribute-binding…

计算机视觉与模式识别 · 计算机科学 2024-05-22 Hong Chen , Xin Wang , Yipeng Zhang , Yuwei Zhou , Zeyang Zhang , Siao Tang , Wenwu Zhu

Image style transfer models based on convolutional neural networks usually suffer from high temporal inconsistency when applied to videos. Some video style transfer models have been proposed to improve temporal consistency, yet they fail to…

计算机视觉与模式识别 · 计算机科学 2018-11-02 Chang Gao , Derun Gu , Fangjun Zhang , Yizhou Yu

The open-domain video generation models are constrained by the scale of the training video datasets, and some less common actions still cannot be generated. Some researchers explore video editing methods and achieve action generation by…

计算机视觉与模式识别 · 计算机科学 2024-08-26 Jianzhi Liu , Junchen Zhu , Lianli Gao , Heng Tao Shen , Jingkuan Song

Current state-of-the-art methods for video inpainting typically rely on optical flow or attention-based approaches to inpaint masked regions by propagating visual information across frames. While such approaches have led to significant…

This paper presents a novel framework termed Cut-and-Paste for real-word semantic video editing under the guidance of text prompt and additional reference image. While the text-driven video editing has demonstrated remarkable ability to…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Zhichao Zuo , Zhao Zhang , Yan Luo , Yang Zhao , Haijun Zhang , Yi Yang , Meng Wang

The conventional training approach for image captioning involves pre-training a network using teacher forcing and subsequent fine-tuning with Self-Critical Sequence Training to maximize hand-crafted captioning metrics. However, when…

计算机视觉与模式识别 · 计算机科学 2024-08-28 Nicholas Moratelli , Davide Caffagni , Marcella Cornia , Lorenzo Baraldi , Rita Cucchiara

In-context image generation and editing (ICGE) enables users to specify visual concepts through interleaved image-text prompts, demanding precise understanding and faithful execution of user intent. Although recent unified multimodal models…

计算机视觉与模式识别 · 计算机科学 2026-01-09 Runze He , Yiji Cheng , Tiankai Hang , Zhimin Li , Yu Xu , Zijin Yin , Shiyi Zhang , Wenxun Dai , Penghui Du , Ao Ma , Chunyu Wang , Qinglin Lu , Jizhong Han , Jiao Dai

In computer vision, it is well-known that a lack of data diversity will impair model performance. In this study, we address the challenges of enhancing the dataset diversity problem in order to benefit various downstream tasks such as…

计算机视觉与模式识别 · 计算机科学 2024-08-02 Yuhang Li , Xin Dong , Chen Chen , Weiming Zhuang , Lingjuan Lyu

This work presents ViGeo, a feed-forward foundation model for recovering spatially dense and temporally consistent geometry from video sequences. Built upon a plain transformer architecture without task-specific architectural modifications,…

计算机视觉与模式识别 · 计算机科学 2026-05-29 Zhu Yu , Jingnan Gao , Runmin Zhang , Lingteng Qiu , Zhengyi Zhao , Rui Peng , Yichao Yan , Kejie Qiu , Siyu Zhu , Si-Yuan Cao , Hui-Liang Shen

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided image diffusion models to image-guided video generation…

计算机视觉与模式识别 · 计算机科学 2024-09-17 Cong Wang , Jiaxi Gu , Panwen Hu , Songcen Xu , Hang Xu , Xiaodan Liang

We propose a novel inference technique based on a pretrained diffusion model for text-conditional video generation. Our approach, called FIFO-Diffusion, is conceptually capable of generating infinitely long videos without additional…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Jihwan Kim , Junoh Kang , Jinyoung Choi , Bohyung Han

High-quality training triplets (instruction, original image, edited image) are essential for instruction-based image editing. Predominant training datasets (e.g., InsPix2Pix) are created using text-to-image generative models (e.g., Stable…

计算机视觉与模式识别 · 计算机科学 2025-03-21 Xin Gu , Ming Li , Libo Zhang , Fan Chen , Longyin Wen , Tiejian Luo , Sijie Zhu

Diffusion models have achieved remarkable advancements in text-to-image generation. However, existing models still have many difficulties when faced with multiple-object compositional generation. In this paper, we propose RealCompo, a new…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Xinchen Zhang , Ling Yang , Yaqi Cai , Zhaochen Yu , Kai-Ni Wang , Jiake Xie , Ye Tian , Minkai Xu , Yong Tang , Yujiu Yang , Bin Cui

In-context generation significantly enhances Diffusion Transformers (DiTs) by enabling controllable image-to-image generation through reference examples. However, the resulting input concatenation drastically increases sequence length,…

计算机视觉与模式识别 · 计算机科学 2026-02-03 Junqing Lin , Xingyu Zheng , Pei Cheng , Bin Fu , Jingwei Sun , Guangzhong Sun

Regression models often fail to generalize effectively in regions characterized by highly imbalanced label distributions. Previous methods for deep imbalanced regression rely on gradient-based weight updates, which tend to overfit in…

机器学习 · 计算机科学 2024-11-21 Ismail Nejjar , Faez Ahmed , Olga Fink

Diffusion models have made significant advances in generating high-quality images, but their application to video generation has remained challenging due to the complexity of temporal motion. Zero-shot video editing offers a solution by…

计算机视觉与模式识别 · 计算机科学 2023-12-20 Xirui Li , Chao Ma , Xiaokang Yang , Ming-Hsuan Yang

Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing…

计算机视觉与模式识别 · 计算机科学 2024-07-30 Yu Lu , Yuanzhi Liang , Linchao Zhu , Yi Yang

Humans do not perceive all parts of a scene with the same resolution, but rather focus on few regions of interest (ROIs). Traditional Object-Based codecs take advantage of this biological intuition, and are capable of non-uniform allocation…

图像与视频处理 · 电气工程与系统科学 2022-11-03 Yura Perugachi-Diaz , Guillaume Sautière , Davide Abati , Yang Yang , Amirhossein Habibian , Taco S Cohen

Recent advances in video reward models and post-training strategies have improved text-to-video (T2V) generation. While these models typically assess visual quality, motion quality, and text alignment, they often overlook key structural…

计算机视觉与模式识别 · 计算机科学 2026-03-27 Yuan Wang , Borui Liao , Huijuan Huang , Jinda Lu , Ouxiang Li , Kuien Liu , Meng Wang , Xiang Wang