中文
相关论文

相关论文: MultiShotMaster: A Controllable Multi-Shot Video G…

200 篇论文

Recent advances in diffusion models bring new vitality to visual content creation. However, current text-to-video generation models still face significant challenges such as high training costs, substantial data requirements, and…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Sicong Feng , Jielong Yang , Li Peng

There has been a recent explosion of impressive generative models that can produce high quality images (or videos) conditioned on text descriptions. However, all such approaches rely on conditional sentences that contain unambiguous…

计算机视觉与模式识别 · 计算机科学 2023-05-09 Tanzila Rahman , Hsin-Ying Lee , Jian Ren , Sergey Tulyakov , Shweta Mahajan , Leonid Sigal

Recent advancements in generative models have revolutionized video synthesis and editing. However, the scarcity of diverse, high-quality datasets continues to hinder video-conditioned robotic learning, limiting cross-platform…

计算机视觉与模式识别 · 计算机科学 2025-10-14 Yang Bai , Liudi Yang , George Eskandar , Fengyi Shen , Dong Chen , Mohammad Altillawi , Ziyuan Liu , Gitta Kutyniok

Despite the rapid progress of video generation models, the role of data in influencing motion is poorly understood. We present Motive (MOTIon attribution for Video gEneration), a motion-centric, gradient-based data attribution framework…

计算机视觉与模式识别 · 计算机科学 2026-01-14 Xindi Wu , Despoina Paschalidou , Jun Gao , Antonio Torralba , Laura Leal-Taixé , Olga Russakovsky , Sanja Fidler , Jonathan Lorraine

Simultaneous processing of multiple video sources requires each pixel in a frame from a video source to be processed synchronously with the pixels at the same spatial positions in corresponding frames from the other video sources. However,…

图像与视频处理 · 电气工程与系统科学 2021-03-02 Mohamed Awad , Islam T. Abougindia , Ahmed Elliethy , Hussein A. Aly

Diffusion models have demonstrated remarkable and robust abilities in both image and video generation. To achieve greater control over generated results, researchers introduce additional architectures, such as ControlNet, Adapters and…

计算机视觉与模式识别 · 计算机科学 2025-03-11 Bohao Peng , Jian Wang , Yuechen Zhang , Wenbo Li , Ming-Chang Yang , Jiaya Jia

The synthesis of spatiotemporally coherent 4D content presents fundamental challenges in computer vision, requiring simultaneous modeling of high-fidelity spatial representations and physically plausible temporal dynamics. Current…

计算机视觉与模式识别 · 计算机科学 2025-12-01 Xiaoyan Liu , Kangrui Li , Yuehao Song , Jiaxin Liu

In this work we present a novel, robust transition generation technique that can serve as a new tool for 3D animators, based on adversarial recurrent neural networks. The system synthesizes high-quality motions that use temporally-sparse…

计算机视觉与模式识别 · 计算机科学 2021-02-10 Félix G. Harvey , Mike Yurick , Derek Nowrouzezahrai , Christopher Pal

This paper presents \emph{ControlVideo} for text-driven video editing -- generating a video that aligns with a given text while preserving the structure of the source video. Building on a pre-trained text-to-image diffusion model,…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Min Zhao , Rongzhen Wang , Fan Bao , Chongxuan Li , Jun Zhu

Large-scale video generation models have the inherent ability to realistically model natural scenes. In this paper, we demonstrate that through a careful design of a generative video propagation framework, various video tasks can be…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Shaoteng Liu , Tianyu Wang , Jui-Hsien Wang , Qing Liu , Zhifei Zhang , Joon-Young Lee , Yijun Li , Bei Yu , Zhe Lin , Soo Ye Kim , Jiaya Jia

In recent years, generative artificial intelligence has achieved significant advancements in the field of image generation, spawning a variety of applications. However, video generation still faces considerable challenges in various…

计算机视觉与模式识别 · 计算机科学 2025-06-30 Yuang Zhang , Jiaxi Gu , Li-Wen Wang , Han Wang , Junqi Cheng , Yuefeng Zhu , Fangyuan Zou

Recent advances in video super-resolution have shown that convolutional neural networks combined with motion compensation are able to merge information from multiple low-resolution (LR) frames to generate high-quality images. Current…

计算机视觉与模式识别 · 计算机科学 2018-03-28 Mehdi S. M. Sajjadi , Raviteja Vemulapalli , Matthew Brown

We present Motion Marionette, a zero-shot framework for rigid motion transfer from monocular source videos to single-view target images. Previous works typically employ geometric, generative, or simulation priors to guide the transfer…

计算机视觉与模式识别 · 计算机科学 2025-11-26 Haoxuan Wang , Jiachen Tao , Junyi Wu , Gaowen Liu , Ramana Rao Kompella , Yan Yan

While text-to-video diffusion models have made significant strides, many still face challenges in generating videos with temporal consistency. Within diffusion frameworks, guidance techniques have proven effective in enhancing output…

计算机视觉与模式识别 · 计算机科学 2025-03-25 Hyelin Nam , Jaemin Kim , Dohun Lee , Jong Chul Ye

Despite advancements in generating visually stunning content, video diffusion models (VDMs) often yield physically inconsistent results due to pixel-only reconstruction. To address this, we propose MMPhysVideo, the first framework to scale…

计算机视觉与模式识别 · 计算机科学 2026-04-06 Shubo Lin , Xuanyang Zhang , Wei Cheng , Weiming Hu , Gang Yu , Jin Gao

Customized text-to-video generation aims to produce high-quality videos that incorporate user-specified subject identities or motion patterns. However, existing methods mainly focus on personalizing a single concept, either subject identity…

计算机视觉与模式识别 · 计算机科学 2025-03-28 Chi-Pin Huang , Yen-Siang Wu , Hung-Kai Chung , Kai-Po Chang , Fu-En Yang , Yu-Chiang Frank Wang

The rise of short-form video platforms and the emergence of multimodal large language models (MLLMs) have amplified the need for scalable, effective, zero-shot text-to-video retrieval systems. While recent advances in large-scale…

信息检索 · 计算机科学 2026-02-24 Jiaxin Wu , Xiao-Yong Wei , Qing Li

While video foundation models excel at single-shot generation, real-world cinematic storytelling inherently relies on complex multi-shot sequencing. Further progress is constrained by the absence of datasets that address three core…

计算机视觉与模式识别 · 计算机科学 2026-05-12 Haojie Zhang , Di Wu , Bingyan Liu , Linjie Zhong , Yuancheng Wei , Xingsong Ye , Nanqing Liu , Yaling Liang

We pose a new problem, In-2-4D, for generative 4D (i.e., 3D + motion) inbetweening to interpolate two single-view images. In contrast to video/4D generation from only text or a single image, our interpolative task can leverage more precise…

图形学 · 计算机科学 2025-09-30 Sauradip Nag , Daniel Cohen-Or , Hao Zhang , Ali Mahdavi-Amiri

In this paper, we introduce VideoNarrator, a novel training-free pipeline designed to generate dense video captions that offer a structured snapshot of video content. These captions offer detailed narrations with precise timestamps,…

计算机视觉与模式识别 · 计算机科学 2025-07-24 Tz-Ying Wu , Tahani Trigui , Sharath Nittur Sridhar , Anand Bodas , Subarna Tripathi