English
Related papers

Related papers: A$^2$RD: Agentic Autoregressive Diffusion for Long…

200 papers

Long-form video understanding remains a fundamental challenge for current Video Large Language Models. Most existing models rely on static reasoning over uniformly sampled frames, which weakens temporal localization and leads to substantial…

Computer Vision and Pattern Recognition · Computer Science 2026-04-21 Chenglin Li , Qianglong Chen , Feng Han , Yikun Wang , Xingxi Yin , Yan Gong , Ruilin Li , Yin Zhang , Jiaqi Wang

Long-form video understanding presents significant challenges due to extensive temporal-spatial complexity and the difficulty of question answering under such extended contexts. While Large Language Models (LLMs) have demonstrated…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Xiaoyi Zhang , Zhaoyang Jia , Zongyu Guo , Jiahao Li , Bin Li , Houqiang Li , Yan Lu

Video-to-video synthesis poses significant challenges in maintaining character consistency, smooth temporal transitions, and preserving visual quality during fast motion. While recent fully cross-frame self-attention mechanisms have…

Computer Vision and Pattern Recognition · Computer Science 2024-11-12 Tanvir Mahmud , Mustafa Munir , Radu Marculescu , Diana Marculescu

Generating realistic human motion sequences from text descriptions is a challenging task that requires capturing the rich expressiveness of both natural language and human motion.Recent advances in diffusion models have enabled significant…

Computer Vision and Pattern Recognition · Computer Science 2023-12-22 Beibei Jing , Youjia Zhang , Zikai Song , Junqing Yu , Wei Yang

Closed-loop driving simulation requires real-time interaction beyond short offline clips, pushing current driving world models toward autoregressive (AR) rollout. Existing AR distillation approaches typically rely on frame sinks or…

Computer Vision and Pattern Recognition · Computer Science 2026-05-25 Conglang Zhang , Yifan Zhan , Qingjie Wang , Zhanpeng Ouyang , Yu Li , Zihao Yang , Xiaoyang Guo , Weiqiang Ren , Qian Zhang , Zhen Dong , Yinqiang Zheng , Wei Yin , Zhengqing Chen

Autoregressive (AR) diffusion offers a promising framework for generating videos of theoretically infinite length. However, a major challenge is maintaining temporal continuity while preventing the progressive quality degradation caused by…

Computer Vision and Pattern Recognition · Computer Science 2026-03-10 Kai Zou , Dian Zheng , Hongbo Liu , Tiankai Hang , Bin Liu , Nenghai Yu

Multimodal generative models have shown remarkable progress in single-modality video and audio synthesis, yet truly joint audio-video generation remains an open challenge. In this paper, I explore four key contributions to advance this…

Sound · Computer Science 2026-03-18 Alejandro Paredes La Torre

Recent advances in video generation can produce realistic, minute-long single-shot videos with scalable diffusion transformers. However, real-world narrative videos require multi-shot scenes with visual and dynamic consistency across shots.…

Computer Vision and Pattern Recognition · Computer Science 2025-03-14 Yuwei Guo , Ceyuan Yang , Ziyan Yang , Zhibei Ma , Zhijie Lin , Zhenheng Yang , Dahua Lin , Lu Jiang

Recently, autoregressive (AR) video diffusion models have achieved remarkable performance. However, due to their limited training durations, a train-test gap emerges when testing at longer horizons, leading to rapid visual degradations.…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Haodong Li , Shaoteng Liu , Zhe Lin , Manmohan Chandraker

Auto-regressive video generation enables long video synthesis by iteratively conditioning each new batch of frames on previously generated content. However, recent work has shown that such pipelines suffer from severe temporal drift, where…

Computer Vision and Pattern Recognition · Computer Science 2026-02-03 Ariel Shaulov , Eitan Shaar , Amit Edenzon , Lior Wolf

Novel architectures have recently improved generative image synthesis leading to excellent visual quality in various tasks. Of particular note is the field of ``AI-Art'', which has seen unprecedented growth with the emergence of powerful…

Computer Vision and Pattern Recognition · Computer Science 2022-07-27 Robin Rombach , Andreas Blattmann , Björn Ommer

Consistency models have demonstrated powerful capability in efficient image generation and allowed synthesis within a few sampling steps, alleviating the high computational cost in diffusion models. However, the consistency model in the…

Computer Vision and Pattern Recognition · Computer Science 2023-12-15 Xiang Wang , Shiwei Zhang , Han Zhang , Yu Liu , Yingya Zhang , Changxin Gao , Nong Sang

Long-form video understanding remains challenging due to the extended temporal structure and dense multimodal cues. Despite recent progress, many existing approaches still rely on hand-crafted reasoning pipelines or employ token-consuming…

Computer Vision and Pattern Recognition · Computer Science 2026-03-31 Yufei Yin , Qianke Meng , Minghao Chen , Jiajun Ding , Zhenwei Shao , Zhou Yu

Building video world models upon pretrained video generation systems represents an important yet challenging step toward general spatiotemporal intelligence. A world model should possess three essential properties: controllability,…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jianxiong Gao , Zhaoxi Chen , Xian Liu , Junhao Zhuang , Chengming Xu , Jianfeng Feng , Yu Qiao , Yanwei Fu , Chenyang Si , Ziwei Liu

Recently, with the tremendous success of diffusion models in the field of text-to-image (T2I) generation, increasing attention has been directed toward their potential in text-to-video (T2V) applications. However, the computational demands…

Computer Vision and Pattern Recognition · Computer Science 2025-01-29 Wenfeng Lin , Jiangchuan Wei , Boyuan Liu , Yichen Zhang , Shiyue Yan , Mingyu Guo

Recent advances in large reconstruction and generative models have significantly improved scene reconstruction and novel view generation. However, due to compute limitations, each inference with these large models is confined to a small…

Computer Vision and Pattern Recognition · Computer Science 2025-04-15 Shangjin Zhai , Zhichao Ye , Jialin Liu , Weijian Xie , Jiaqi Hu , Zhen Peng , Hua Xue , Danpeng Chen , Xiaomeng Wang , Lei Yang , Nan Wang , Haomin Liu , Guofeng Zhang

Video diffusion models provide powerful real-world simulators for embodied AI but remain limited in controllability for robotic manipulation. Recent works on trajectory-conditioned video generation address this gap but often rely on 2D…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Yang Bai , Liudi Yang , George Eskandar , Fengyi Shen , Mohammad Altillawi , Ziyuan Liu , Gitta Kutyniok

Video diffusion models have made substantial progress in various video generation applications. However, training models for long video generation tasks require significant computational and data resources, posing a challenge to developing…

Computer Vision and Pattern Recognition · Computer Science 2024-07-30 Yu Lu , Yuanzhi Liang , Linchao Zhu , Yi Yang

Video enhancement is a challenging problem, more than that of stills, mainly due to high computational cost, larger data volumes and the difficulty of achieving consistency in the spatio-temporal domain. In practice, these challenges are…

Image and Video Processing · Electrical Eng. & Systems 2022-12-13 Dario Fuoli , Zhiwu Huang , Danda Pani Paudel , Luc Van Gool , Radu Timofte

Retrieval-augmented generation (RAG) has emerged as a pivotal method for expanding the knowledge of large language models. To handle complex queries more effectively, researchers developed Adaptive-RAG (A-RAG) to enhance the generated…

Artificial Intelligence · Computer Science 2025-05-27 Jie Ou , Jinyu Guo , Shuaihong Jiang , Zhaokun Wang , Libo Qin , Shunyu Yao , Wenhong Tian
‹ Prev 1 8 9 10 Next ›