中文
相关论文

相关论文: MoCA-Video: Motion-Aware Concept Alignment for Con…

200 篇论文

Image generation and editing have seen a great deal of advancements with the rise of large-scale diffusion models that allow user control of different modalities such as text, mask, depth maps, etc. However, controlled editing of videos…

计算机视觉与模式识别 · 计算机科学 2024-06-04 AmirHossein Zamani , Amir G. Aghdam , Tiberiu Popa , Eugene Belilovsky

In this report, we present MagicEdit, a surprisingly simple yet effective solution to the text-guided video editing task. We found that high-fidelity and temporally coherent video-to-video translation can be achieved by explicitly…

计算机视觉与模式识别 · 计算机科学 2023-08-29 Jun Hao Liew , Hanshu Yan , Jianfeng Zhang , Zhongcong Xu , Jiashi Feng

Text-to-image diffusion model alignment is critical for improving the alignment between the generated images and human preferences. While training-based methods are constrained by high computational costs and dataset requirements,…

计算机视觉与模式识别 · 计算机科学 2025-03-26 Xin Xie , Dong Gong

Given an untrimmed video and a language query depicting a specific temporal moment in the video, video grounding aims to localize the time interval by understanding the text and video simultaneously. One of the most challenging issues is an…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Dahye Kim , Jungin Park , Jiyoung Lee , Seongheon Park , Kwanghoon Sohn

Generative novel view synthesis faces a fundamental dilemma: geometric priors provide spatial alignment but become sparse and inaccurate under view changes, while appearance priors offer visual fidelity but lack geometric correspondence.…

计算机视觉与模式识别 · 计算机科学 2026-05-14 Haofeng Liu , Yang Zhou , Ziheng Wang , Zhengbo Xu , Zhan Peng , Jie Ma , Jun Liang , Shengfeng He , Jing Li

Dense video captioning jointly localizes and captions salient events in untrimmed videos. Recent methods primarily focus on leveraging additional prior knowledge and advanced multi-task architectures to achieve competitive performance.…

计算机视觉与模式识别 · 计算机科学 2025-11-14 Mingda Jia , Weiliang Meng , Zenghuang Fu , Yiheng Li , Qi Zeng , Yifan Zhang , Ju Xin , Rongtao Xu , Jiguang Zhang , Xiaopeng Zhang

Temporally consistent dense video annotations are scarce and hard to collect. In contrast, image segmentation datasets (and pre-trained models) are ubiquitous, and easier to label for any novel task. In this paper, we introduce a method to…

计算机视觉与模式识别 · 计算机科学 2022-03-18 Aharon Azulay , Tavi Halperin , Orestis Vantzos , Nadav Borenstein , Ofir Bibi

Human beings have the ability to continuously analyze a video and immediately extract the motion components. We want to adopt this paradigm to provide a coherent and stable motion segmentation over the video sequence. In this perspective,…

计算机视觉与模式识别 · 计算机科学 2024-04-18 Etienne Meunier , Patrick Bouthemy

Text-conditioned image editing has succeeded in various types of editing based on a diffusion framework. Unfortunately, this success did not carry over to a video, which continues to be challenging. Existing video editing systems are still…

计算机视觉与模式识别 · 计算机科学 2023-12-13 Sunjae Yoon , Gwanhyeong Koo , Ji Woo Hong , Chang D. Yoo

We present a method to generate video-action pairs that follow text instructions, starting from an initial image observation and the robot's joint states. Our approach automatically provides action labels for video diffusion models,…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Liudi Yang , Yang Bai , George Eskandar , Fengyi Shen , Mohammad Altillawi , Dong Chen , Ziyuan Liu , Abhinav Valada

Continuous image editing aims to provide slider-style control of edit strength while preserving source-image fidelity and maintaining a consistent edit direction. Existing learning-based slider methods typically rely on auxiliary modules…

计算机视觉与模式识别 · 计算机科学 2026-04-03 Taichi Endo , Guoqing Hao , Kazuhiko Sumi

Recent years have seen a tremendous improvement in the quality of video generation and editing approaches. While several techniques focus on editing appearance, few address motion. Current approaches using text, trajectories, or bounding…

计算机视觉与模式识别 · 计算机科学 2025-05-27 Manuel Kansy , Jacek Naruniec , Christopher Schroers , Markus Gross , Romann M. Weber

Recent advancements in Large Language Models have successfully transitioned towards System 2 reasoning, yet applying these paradigms to video understanding remains challenging. While prevailing research attributes failures in Video-LLMs to…

计算机视觉与模式识别 · 计算机科学 2026-01-08 Hongbo Jin , Jiayu Ding , Siyi Xie , Guibo Luo , Ge Li

Text-driven video editing aims to modify video content based on natural language instructions. While recent training-free methods have leveraged pretrained diffusion models, they often rely on an inversion-editing paradigm. This paradigm…

计算机视觉与模式识别 · 计算机科学 2025-12-15 Guangzhao Li , Yanming Yang , Chenxi Song , Chi Zhang

Instruction-based video editing aims to modify an input video according to a natural-language instruction while preserving content fidelity and temporal coherence. However, existing diffusion-based approaches are often trained on paired…

计算机视觉与模式识别 · 计算机科学 2025-12-19 Xiaoyan Cong , Haotian Yang , Angtian Wang , Yizhi Wang , Yiding Yang , Canyu Zhang , Chongyang Ma

Recent works in video prediction have mainly focused on passive forecasting and low-level action-conditional prediction, which sidesteps the learning of interaction between agents and objects. We introduce the task of semantic…

计算机视觉与模式识别 · 计算机科学 2022-04-27 Wei Yu , Wenxin Chen , Songhenh Yin , Steve Easterbrook , Animesh Garg

For semantic segmentation, most existing real-time deep models trained with each frame independently may produce inconsistent results for a video sequence. Advanced methods take into considerations the correlations in the video sequence,…

计算机视觉与模式识别 · 计算机科学 2020-07-20 Yifan Liu , Chunhua Shen , Changqian Yu , Jingdong Wang

Editing images with diffusion models under strict training-free constraints remains a significant challenge. While recent optimisation-based methods achieve strong zero-shot edits from text, they struggle to preserve identity and capture…

计算机视觉与模式识别 · 计算机科学 2026-03-04 Niki Foteinopoulou , Ignas Budvytis , Stephan Liwicki

Weakly-supervised audio-visual video parsing (AVVP) seeks to detect audible, visible, and audio-visual events without temporal annotations. Previous work has emphasized refining global predictions through contrastive or collaborative…

计算机视觉与模式识别 · 计算机科学 2025-09-18 Yaru Chen , Ruohao Guo , Liting Gao , Yang Xiang , Qingyu Luo , Zhenbo Li , Wenwu Wang

Temporal consistency is essential for video editing applications. Existing work on layered representation of videos allows propagating edits consistently to each frame. These methods, however, can only edit object appearance rather than…

计算机视觉与模式识别 · 计算机科学 2023-01-31 Yao-Chih Lee , Ji-Ze Genevieve Jang , Yi-Ting Chen , Elizabeth Qiu , Jia-Bin Huang