English
Related papers

Related papers: MAGI-1: Autoregressive Video Generation at Scale

200 papers

Recent interactive video world model methods generate scene evolution conditioned on user instructions. Although they achieve impressive results, two key limitations remain. First, they exhibit motion drift in complex environments with…

Computer Vision and Pattern Recognition · Computer Science 2026-03-19 Guangyuan Li , Bo Li , Jinwei Chen , Xiaobin Hu , Lei Zhao , Peng-Tao Jiang

Controllable ultra-long video generation is a fundamental yet challenging task. Although existing methods are effective for short clips, they struggle to scale due to issues such as temporal inconsistency and visual degradation. In this…

Computer Vision and Pattern Recognition · Computer Science 2025-08-06 Jianxiong Gao , Zhaoxi Chen , Xian Liu , Jianfeng Feng , Chenyang Si , Yanwei Fu , Yu Qiao , Ziwei Liu

Diffusion based video generation has received extensive attention and achieved considerable success within both the academic and industrial communities. However, current efforts are mainly concentrated on single-objective or single-task…

Computer Vision and Pattern Recognition · Computer Science 2024-01-18 Ludan Ruan , Lei Tian , Chuanwei Huang , Xu Zhang , Xinyan Xiao

Predicting human gaze in video is fundamental to advancing scene understanding and multimodal interaction. While traditional saliency maps provide spatial probability distributions and scanpaths offer ordered fixations, both abstractions…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Jenna Kang , Colin Groth , Tong Wu , Finley Torrens , Patsorn Sangkloy , Gordon Wetzstein , Qi Sun

We introduce layered controllable video generation, where we, without any supervision, decompose the initial frame of a video into foreground and background layers, with which the user can control the video generation process by simply…

Computer Vision and Pattern Recognition · Computer Science 2022-10-05 Jiahui Huang , Yuhe Jin , Kwang Moo Yi , Leonid Sigal

We present Waver, a high-performance foundation model for unified image and video generation. Waver can directly generate videos with durations ranging from 5 to 10 seconds at a native resolution of 720p, which are subsequently upscaled to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Yifu Zhang , Hao Yang , Yuqi Zhang , Yifei Hu , Fengda Zhu , Chuang Lin , Xiaofeng Mei , Yi Jiang , Bingyue Peng , Zehuan Yuan

Video generation has many unique challenges beyond those of image generation. The temporal dimension introduces extensive possible variations across frames, over which consistency and continuity may be violated. In this study, we move…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Weixi Feng , Jiachen Li , Michael Saxon , Tsu-jui Fu , Wenhu Chen , William Yang Wang

Despite the recent strides in video generation, state-of-the-art methods still struggle with elements of visual detail. One particularly challenging case is the class of videos in which the intricate motion of the hand coupled with a mostly…

Computer Vision and Pattern Recognition · Computer Science 2025-07-15 Yayuan Li , Zhi Cao , Jason J. Corso

Current video generation models perform well at single-shot synthesis but struggle with multi-shot videos, facing critical challenges in maintaining character and background consistency across shots and flexibly generating videos of…

Computer Vision and Pattern Recognition · Computer Science 2025-12-15 Xiangyang Luo , Qingyu Li , Xiaokun Liu , Wenyu Qin , Miao Yang , Meng Wang , Pengfei Wan , Di Zhang , Kun Gai , Shao-Lun Huang

We introduce Motion-I2V, a novel framework for consistent and controllable image-to-video generation (I2V). In contrast to previous methods that directly learn the complicated image-to-video mapping, Motion-I2V factorizes I2V into two…

Computer Vision and Pattern Recognition · Computer Science 2024-02-01 Xiaoyu Shi , Zhaoyang Huang , Fu-Yun Wang , Weikang Bian , Dasong Li , Yi Zhang , Manyuan Zhang , Ka Chun Cheung , Simon See , Hongwei Qin , Jifeng Dai , Hongsheng Li

Generating long and consistent videos has emerged as a significant yet challenging problem. While most existing diffusion-based video generation models, derived from image generation models, demonstrate promising performance in generating…

Computer Vision and Pattern Recognition · Computer Science 2024-04-30 Yichen Ouyang , jianhao Yuan , Hao Zhao , Gaoang Wang , Bo zhao

The rapid advancement of video generation models has enabled the creation of highly realistic synthetic media, raising significant societal concerns regarding the spread of misinformation. However, current detection methods suffer from…

Computer Vision and Pattern Recognition · Computer Science 2026-04-07 Zhengcen Li , Chenyang Jiang , Hang Zhao , Shiyang Zhou , Yunyang Mo , Feng Gao , Fan Yang , Qiben Shan , Shaocong Wu , Jingyong Su

Interactive humanoid video generation aims to synthesize lifelike visual agents that can engage with humans through continuous and responsive video. Despite recent advances in video synthesis, existing methods often grapple with the…

Large-scale pre-trained video diffusion models have exhibited remarkable capabilities in diverse video generation. However, existing solutions face several challenges in generating long videos with rich human-scene interactions (HSI),…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Zekun Li , Rui Zhou , Rahul Sajnani , Xiaoyan Cong , Daniel Ritchie , Srinath Sridhar

We present MAGNET (Model Autonomously Growing Network), a decentralized system for autonomous generation, training, and serving of domain-expert language models across commodity hardware. MAGNET integrates four components: (1) autoresearch,…

Machine Learning · Computer Science 2026-03-30 Yongwan Kim , Sungchul Park

Generating music that aligns with the visual content of a video has been a challenging task, as it requires a deep understanding of visual semantics and involves generating music whose melody, rhythm, and dynamics harmonize with the visual…

Sound · Computer Science 2024-10-18 Ruiqi Li , Siqi Zheng , Xize Cheng , Ziang Zhang , Shengpeng Ji , Zhou Zhao

We present xGen-VideoSyn-1, a text-to-video (T2V) generation model capable of producing realistic scenes from textual descriptions. Building on recent advancements, such as OpenAI's Sora, we explore the latent diffusion model (LDM)…

Humans excel at forecasting the future dynamics of a scene given just a single image. Video generation models that can mimic this ability are an essential component for intelligent systems. Recent approaches have improved temporal coherence…

Computer Vision and Pattern Recognition · Computer Science 2026-05-18 Melonie de Almeida , Daniela Ivanova , Tong Shi , John H. Williamson , Paul Henderson

The generative model has made significant advancements in the creation of realistic videos, which causes security issues. However, this emerging risk has not been adequately addressed due to the absence of a benchmark dataset for…

Computer Vision and Pattern Recognition · Computer Science 2024-05-08 Peisong He , Leyao Zhu , Jiaxing Li , Shiqi Wang , Haoliang Li

Video generation models, as one form of world models, have emerged as one of the most exciting frontiers in AI, promising agents the ability to imagine the future by modeling the temporal evolution of complex scenes. In autonomous driving,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Yang Zhou , Hao Shao , Letian Wang , Zhuofan Zong , Hongsheng Li , Steven L. Waslander
‹ Prev 1 4 5 6 7 8 10 Next ›