English
Related papers

Related papers: SPIRAL: Self-Evolving Action-Conditioned Video Gen…

200 papers

We present GLASS, a method for Global and Local Action-driven Sequence Synthesis. GLASS is a generative model that is trained on video sequences in an unsupervised manner and that can animate an input image at test time. The method learns…

Computer Vision and Pattern Recognition · Computer Science 2022-04-14 Aram Davtyan , Paolo Favaro

We tackle the long video generation problem, i.e.~generating videos beyond the output length of video generation models. Due to the computation resource constraints, video generation models can only generate video clips that are relatively…

Computer Vision and Pattern Recognition · Computer Science 2023-04-18 Hsin-Ping Huang , Yu-Chuan Su , Ming-Hsuan Yang

Real-world videos consist of sequences of events. Generating such sequences with precise temporal control is infeasible with existing video generators that rely on a single paragraph of text as input. When tasked with generating multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Ziyi Wu , Aliaksandr Siarohin , Willi Menapace , Ivan Skorokhodov , Yuwei Fang , Varnith Chordia , Igor Gilitschenski , Sergey Tulyakov

We present visual action prompts, a unified action representation for action-to-video generation of complex high-DoF interactions while maintaining transferable visual dynamics across domains. Action-driven video generation faces a…

Computer Vision and Pattern Recognition · Computer Science 2025-08-19 Yuang Wang , Chao Wen , Haoyu Guo , Sida Peng , Minghan Qin , Hujun Bao , Xiaowei Zhou , Ruizhen Hu

This paper presents a new method of engaging older participants in the process of application and IT solutions development for older adults for emerging IT and tech startups. A new method called SPIRAL (Support for Participant Involvement…

Software Engineering · Computer Science 2018-03-28 Wiesław Kopeć , Radosław Nielek , Adam Wierzbicki

Recent advancements in personalized Text-to-Video (T2V) generation have made significant strides in synthesizing character-specific content. However, these methods face a critical limitation: the inability to perform fine-grained control…

Computer Vision and Pattern Recognition · Computer Science 2026-01-06 Haopeng Fang , Di Qiu , Binjie Mao , He Tang

Custom Storyboard Generation (CSG) aims to produce high-quality, multi-character consistent storytelling. Current approaches based on static diffusion models, whether used in a one-shot manner or within multi-agent frameworks, face three…

Computer Vision and Pattern Recognition · Computer Science 2026-02-25 Hailong Yan , Shice Liu , Tao Wang , Xiangtao Zhang , Yijie Zhong , Jinwei Chen , Le Zhang , Bo Li

Video generation models trained on heterogeneous data with likelihood-surrogate objectives can produce visually plausible rollouts that violate physical constraints in embodied manipulation. Although reinforcement-learning post-training…

Computer Vision and Pattern Recognition · Computer Science 2026-05-15 Zhenyang Ni , Yijiang Li , Ruochen Jiao , Simon Sinong Zhan , Sipeng Chen , Zhenfei Yin , Minshuo Chen , Philip Torr , Zhaoran Wang , Qi Zhu

Audio-driven video generation aims to synthesize realistic videos that align with input audio recordings, akin to the human ability to visualize scenes from auditory input. However, existing approaches predominantly focus on exploring…

Graphics · Computer Science 2026-03-17 Kien T. Pham , Yingqing He , Yazhou Xing , Qifeng Chen , Long Chen

Advances in deep generative networks have led to impressive results in recent years. Nevertheless, such models can often waste their capacity on the minutiae of datasets, presumably due to weak inductive biases in their decoders. This is…

Computer Vision and Pattern Recognition · Computer Science 2018-04-05 Yaroslav Ganin , Tejas Kulkarni , Igor Babuschkin , S. M. Ali Eslami , Oriol Vinyals

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Lijie Liu , Tianxiang Ma , Bingchuan Li , Zhuowei Chen , Jiawei Liu , Gen Li , Siyu Zhou , Qian He , Xinglong Wu

Recent advances in text-to-video generation have achieved impressive performance on short clips, yet evaluating long-form generation under complex textual inputs remains a significant challenge. In response to this challenge, we present…

Computer Vision and Pattern Recognition · Computer Science 2026-05-29 Xiangqing Zheng , Chengyue Wu , Kehai Chen , Min Zhang

The "one-shot" technique represents a distinct and sophisticated aesthetic in filmmaking. However, its practical realization is often hindered by prohibitive costs and complex real-world constraints. Although emerging video generation…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Jiawei Liu , Junqiao Li , Jiangfan Deng , Gen Li , Siyu Zhou , Zetao Fang , Shanshan Lao , Zengde Deng , Jianing Zhu , Tingting Ma , Jiayi Li , Yunqiu Wang , Qian He , Xinglong Wu

Recent advances in audio-driven avatar video generation have significantly enhanced audio-visual realism. However, existing methods treat instruction conditioning merely as low-level tracking driven by acoustic or visual cues, without…

Computer Vision and Pattern Recognition · Computer Science 2025-09-18 Yikang Ding , Jiwen Liu , Wenyuan Zhang , Zekun Wang , Wentao Hu , Liyuan Cui , Mingming Lao , Yingchao Shao , Hui Liu , Xiaohan Li , Ming Chen , Xiaoqiang Liu , Yu-Shen Liu , Pengfei Wan

Recent large vision-language models have achieved strong performance on short- and medium-length video understanding, yet they remain inadequate for ultra-long or even infinite video reasoning, where models must preserve coherent memory…

Artificial Intelligence · Computer Science 2026-05-08 Peizheng Yan , Yu Zhao , Liang Xie , Juntong Qi , Mingming Wang , Erwei Yin

Video generation remains a challenging task due to spatiotemporal complexity and the requirement of synthesizing diverse motions with temporal consistency. Previous works attempt to generate videos in arbitrary lengths either in an…

Computer Vision and Pattern Recognition · Computer Science 2023-04-07 Xiaoqian Shen , Xiang Li , Mohamed Elhoseiny

The goal of conditional image-to-video (cI2V) generation is to create a believable new video by beginning with the condition, i.e., one image and text.The previous cI2V generation methods conventionally perform in RGB pixel space, with…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Cuifeng Shen , Yulu Gan , Chen Chen , Xiongwei Zhu , Lele Cheng , Tingting Gao , Jinzhi Wang

To make effective decisions in novel environments with long-horizon goals, it is crucial to engage in hierarchical reasoning across spatial and temporal scales. This entails planning abstract subgoal sequences, visually reasoning about the…

Diffusion Transformers have demonstrated remarkable capabilities in visual synthesis, yet they often struggle with high-level semantic reasoning and long-horizon planning. This limitation frequently leads to visual hallucinations and…

Computer Vision and Pattern Recognition · Computer Science 2025-11-25 Lun Huang , You Xie , Hongyi Xu , Tianpei Gu , Chenxu Zhang , Guoxian Song , Zenan Li , Xiaochen Zhao , Linjie Luo , Guillermo Sapiro

Long-horizon multimodal agents in open-world games must stay goal-directed across many low-level interactions under tight token and latency budgets. Existing approaches often trade off costly per-step reasoning against reactive execution…

Computer Vision and Pattern Recognition · Computer Science 2026-05-19 Wencan Jiang , Jiangning Zhang , Jianbiao Mei , Jinzhuo Liu , Yu Yang , Xiaobin Hu , Zhucun Xue , Yong Liu , Dacheng Tao