中文
相关论文

相关论文: RACCooN: A Versatile Instructional Video Editing F…

200 篇论文

Visual signals in a video can be divided into content and motion. While content specifies which objects are in the video, motion describes their dynamics. Based on this prior, we propose the Motion and Content decomposed Generative…

计算机视觉与模式识别 · 计算机科学 2017-12-15 Sergey Tulyakov , Ming-Yu Liu , Xiaodong Yang , Jan Kautz

We propose a new task, dataset and model for grounded video caption generation. This task unifies captioning and object grounding in video, where the objects in the caption are grounded in the video via temporally consistent bounding boxes.…

计算机视觉与模式识别 · 计算机科学 2024-11-13 Evangelos Kazakos , Cordelia Schmid , Josef Sivic

Significant progress has been made in text-to-video generation through the use of powerful generative models and large-scale internet data. However, substantial challenges remain in precisely controlling individual concepts within the…

计算机视觉与模式识别 · 计算机科学 2024-09-30 Hanxin Zhu , Tianyu He , Anni Tang , Junliang Guo , Zhibo Chen , Jiang Bian

The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual quality and user control over the generated content. In this work, we…

计算机视觉与模式识别 · 计算机科学 2023-11-21 Michal Geyer , Omer Bar-Tal , Shai Bagon , Tali Dekel

While recent text-to-video models excel at generating diverse scenes, they struggle with precise motion control, particularly for complex, multi-subject motions. Although methods for single-motion customization have been developed to…

计算机视觉与模式识别 · 计算机科学 2025-10-28 Youcan Xu , Zhen Wang , Jiaxin Shi , Kexin Li , Feifei Shao , Jun Xiao , Yi Yang , Jun Yu , Long Chen

Text-to-video generation has advanced rapidly, but existing methods typically output only the final composited video and lack editable layered representations, limiting their use in professional workflows. We propose \textbf{LayerT2V}, a…

计算机视觉与模式识别 · 计算机科学 2026-02-27 Guangzhao Li , Kangrui Cen , Baixuan Zhao , Yi Xin , Siqi Luo , Guangtao Zhai , Lei Zhang , Xiaohong Liu

Recent advances in video diffusion models shows promise for generating robotic decision-making data, with trajectory conditions further enabling fine-grained control. However, existing methods primarily focus on individual object motion and…

计算机视觉与模式识别 · 计算机科学 2026-01-28 Xiao Fu , Xintao Wang , Xian Liu , Jianhong Bai , Runsen Xu , Pengfei Wan , Di Zhang , Dahua Lin

Recently, large-scale text-to-image (T2I) models have shown impressive performance in generating high-fidelity images, but with limited controllability, e.g., precisely specifying the content in a specific region with a free-form text…

计算机视觉与模式识别 · 计算机科学 2022-11-29 Zhengyuan Yang , Jianfeng Wang , Zhe Gan , Linjie Li , Kevin Lin , Chenfei Wu , Nan Duan , Zicheng Liu , Ce Liu , Michael Zeng , Lijuan Wang

Causal autoregressive video diffusion models support real-time streaming generation by extrapolating future chunks from previously generated content. Distilling such generators from high-fidelity bidirectional teachers yields competitive…

计算机视觉与模式识别 · 计算机科学 2026-05-15 Yanzuo Lu , Ronglai Zuo , Jiankang Deng

Automatic colorization of line drawings has been widely studied to reduce the labor cost of hand-drawn anime production. Deep learning approaches, including image/video generation and feature-based correspondence, have improved accuracy but…

计算机视觉与模式识别 · 计算机科学 2025-10-02 Kazuma Nagata , Naoshi Kaneko

Effectively handling temporal redundancy remains a key challenge in learning video models. Prevailing approaches often treat each set of frames independently, failing to effectively capture the temporal dependencies and redundancies…

计算机视觉与模式识别 · 计算机科学 2025-07-04 Xiang Fan , Xiaohang Sun , Kushan Thakkar , Zhu Liu , Vimal Bhat , Ranjay Krishna , Xiang Hao

In the dynamic field of digital content creation using generative models, state-of-the-art video editing models still do not offer the level of quality and control that users desire. Previous works on video editing either extended from…

计算机视觉与模式识别 · 计算机科学 2024-11-05 Max Ku , Cong Wei , Weiming Ren , Harry Yang , Wenhu Chen

Video Paragraph Captioning (VPC) aims to generate paragraph captions that summarises key events within a video. Despite recent advancements, challenges persist, notably in effectively utilising multimodal signals inherent in videos and…

计算机视觉与模式识别 · 计算机科学 2024-10-15 Eileen Wang , Caren Han , Josiah Poon

We introduce $\textit{InteractiveVideo}$, a user-centric framework for video generation. Different from traditional generative approaches that operate based on user-provided images or text, our framework is designed for dynamic interaction,…

计算机视觉与模式识别 · 计算机科学 2024-02-06 Yiyuan Zhang , Yuhao Kang , Zhixin Zhang , Xiaohan Ding , Sanyuan Zhao , Xiangyu Yue

Verbal videos, featuring voice-overs or text overlays, provide valuable content but present significant challenges in composition, especially when incorporating editing effects to enhance clarity and visual appeal. In this paper, we…

计算机视觉与模式识别 · 计算机科学 2024-07-08 Weibo Gong , Xiaojie Jin , Xin Li , Dongliang He , Xinglong Wu

We present an approach that exploits hierarchical Recurrent Neural Networks (RNNs) to tackle the video captioning problem, i.e., generating one or multiple sentences to describe a realistic video. Our hierarchical framework contains a…

计算机视觉与模式识别 · 计算机科学 2016-04-07 Haonan Yu , Jiang Wang , Zhiheng Huang , Yi Yang , Wei Xu

Audio to Video generation is an interesting problem that has numerous applications across industry verticals including film making, multi-media, marketing, education and others. High-quality video generation with expressive facial movements…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Neeraj Kumar , Srishti Goel , Ankur Narang , Mujtaba Hasan

Accurately preserving motion while editing a subject remains a core challenge in video editing tasks. Existing methods often face a trade-off between edit and motion fidelity, as they rely on motion representations that are either…

计算机视觉与模式识别 · 计算机科学 2025-11-25 Yeji Song , Jaehyun Lee , Mijin Koo , JunHoo Lee , Nojun Kwak

We present DreamToNav, a novel autonomous robot framework that uses generative video models to enable intuitive, human-in-the-loop control. Instead of relying on rigid waypoint navigation, users provide natural language prompts (e.g.…

机器人学 · 计算机科学 2026-03-09 Valerii Serpiva , Jeffrin Sam , Chidera Simon , Hajira Amjad , Iana Zhura , Artem Lykov , Dzmitry Tsetserukou

Convolutional video models have an order of magnitude larger computational complexity than their counterpart image-level models. Constrained by computational resources, there is no model or training method that can train long video…

计算机视觉与模式识别 · 计算机科学 2021-03-23 Bo Pang , Gao Peng , Yizhuo Li , Cewu Lu