English
Related papers

Related papers: HunyuanVideo: A Systematic Framework For Large Vid…

200 papers

Training strong video generation models usually requires massive datasets, large parameter counts, and substantial compute. In this work, we ask whether strong text-to-video quality is possible at a much smaller budget: fewer than 10M clips…

We present Wan-Move, a simple and scalable framework that brings motion control to video generative models. Existing motion-controllable methods typically suffer from coarse control granularity and limited scalability, leaving their outputs…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Ruihang Chu , Yefei He , Zhekai Chen , Shiwei Zhang , Xiaogang Xu , Bin Xia , Dingdong Wang , Hongwei Yi , Xihui Liu , Hengshuang Zhao , Yu Liu , Yingya Zhang , Yujiu Yang

Learning to represent and generate videos from unlabeled data is a very challenging problem. To generate realistic videos, it is important not only to ensure that the appearance of each frame is real, but also to ensure the plausibility of…

Computer Vision and Pattern Recognition · Computer Science 2017-12-04 Katsunori Ohnishi , Shohei Yamamoto , Yoshitaka Ushiku , Tatsuya Harada

Video generation models have demonstrated remarkable performance, yet their broader adoption remains constrained by slow inference speeds and substantial computational costs, primarily due to the iterative nature of the denoising process.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Xin Zhou , Dingkang Liang , Kaijin Chen , Tianrui Feng , Xiwu Chen , Hongkai Lin , Yikang Ding , Feiyang Tan , Hengshuang Zhao , Xiang Bai

Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Team Seedance , De Chen , Liyang Chen , Xin Chen , Ying Chen , Zhuo Chen , Zhuowei Chen , Feng Cheng , Tianheng Cheng , Yufeng Cheng , Mojie Chi , Xuyan Chi , Jian Cong , Qinpeng Cui , Fei Ding , Qide Dong , Yujiao Du , Haojie Duanmu , Junliang Fan , Jiarui Fang , Jing Fang , Zetao Fang , Chengjian Feng , Yu Gao , Diandian Gu , Dong Guo , Hanzhong Guo , Qiushan Guo , Boyang Hao , Hongxiang Hao , Haoxun He , Jiaao He , Qian He , Tuyen Hoang , Heng Hu , Ruoqing Hu , Yuxiang Hu , Jiancheng Huang , Weilin Huang , Zhaoyang Huang , Zhongyi Huang , Jishuo Jin , Ming Jing , Ashley Kim , Shanshan Lao , Yichong Leng , Bingchuan Li , Gen Li , Haifeng Li , Huixia Li , Jiashi Li , Ming Li , Xiaojie Li , Xingxing Li , Yameng Li , Yiying Li , Yu Li , Yueyan Li , Chao Liang , Han Liang , Jianzhong Liang , Ying Liang , Wang Liao , J. H. Lien , Shanchuan Lin , Xi Lin , Feng Ling , Yue Ling , Fangfang Liu , Jiawei Liu , Jihao Liu , Jingtuo Liu , Shu Liu , Sichao Liu , Wei Liu , Xue Liu , Zuxi Liu , Ruijie Lu , Lecheng Lyu , Jingting Ma , Tianxiang Ma , Xiaonan Nie , Jingzhe Ning , Junjie Pan , Xitong Pan , Ronggui Peng , Xueqiong Qu , Yuxi Ren , Yuchen Shen , Guang Shi , Lei Shi , Yinglong Song , Fan Sun , Li Sun , Renfei Sun , Wenjing Tang , Boyang Tao , Zirui Tao , Dongliang Wang , Feng Wang , Hulin Wang , Ke Wang , Qingyi Wang , Rui Wang , Shuai Wang , Shulei Wang , Weichen Wang , Xuanda Wang , Yanhui Wang , Yue Wang , Yuping Wang , Yuxuan Wang , Zijie Wang , Ziyu Wang , Guoqiang Wei , Meng Wei , Di Wu , Guohong Wu , Hanjie Wu , Huachao Wu , Jian Wu , Jie Wu , Ruolan Wu , Shaojin Wu , Xiaohu Wu , Xinglong Wu , Yonghui Wu , Ruiqi Xia , Xin Xia , Xuefeng Xiao , Shuang Xu , Bangbang Yang , Jiaqi Yang , Runkai Yang , Tao Yang , Yihang Yang , Zhixian Yang , Ziyan Yang , Fulong Ye , Bingqian Yi , Xing Yin , Yongbin You , Linxiao Yuan , Weihong Zeng , Xuejiao Zeng , Yan Zeng , Siyu Zhai , Zhonghua Zhai , Bowen Zhang , Chenlin Zhang , Heng Zhang , Jun Zhang , Manlin Zhang , Peiyuan Zhang , Shuo Zhang , Xiaohe Zhang , Xiaoying Zhang , Xinyan Zhang , Xinyi Zhang , Yichi Zhang , Zixiang Zhang , Haiyu Zhao , Huating Zhao , Liming Zhao , Yian Zhao , Guangcong Zheng , Jianbin Zheng , Xiaozheng Zheng , Zerong Zheng , Kuan Zhu , Feilong Zuo

Cascaded video super-resolution has emerged as a promising technique for decoupling the computational burden associated with generating high-resolution videos using large foundation models. Existing studies, however, are largely confined to…

Computer Vision and Pattern Recognition · Computer Science 2025-10-10 Shian Du , Menghan Xia , Chang Liu , Quande Liu , Xintao Wang , Pengfei Wan , Xiangyang Ji

Videos are created to express emotion, exchange information, and share experiences. Video synthesis has intrigued researchers for a long time. Despite the rapid progress driven by advances in visual synthesis, most existing studies focus on…

Computer Vision and Pattern Recognition · Computer Science 2022-09-27 Songwei Ge , Thomas Hayes , Harry Yang , Xi Yin , Guan Pang , David Jacobs , Jia-Bin Huang , Devi Parikh

Humanoid robots, with their human-like form, are uniquely suited for interacting in environments built for people. However, enabling humanoids to reason, plan, and act in complex open-world settings remains a challenge. World models, models…

Robotics · Computer Science 2025-07-10 Muhammad Qasim Ali , Aditya Sridhar , Shahbuland Matiana , Alex Wong , Mohammad Al-Sharman

Video generation has achieved remarkable progress, with generated videos increasingly resembling real ones. However, the rapid advance in generation has outpaced the development of adequate evaluation metrics. Currently, the assessment of…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Nabyl Quignon , Baptiste Chopin , Yaohui Wang , Antitza Dantcheva

Text-based diffusion models have exhibited remarkable success in generation and editing, showing great promise for enhancing visual content with their generative prior. However, applying these models to video super-resolution remains…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Shangchen Zhou , Peiqing Yang , Jianyi Wang , Yihang Luo , Chen Change Loy

Recently developed methods for video analysis, especially models for pose estimation and behavior classification, are transforming behavioral quantification to be more precise, scalable, and reproducible in fields such as neuroscience and…

Quantitative Methods · Quantitative Biology 2023-03-10 Kevin Luxem , Jennifer J. Sun , Sean P. Bradley , Keerthi Krishnan , Eric A. Yttri , Jan Zimmermann , Talmo D. Pereira , Mark Laubach

OpenKinoAI is an open source framework for post-production of ultra high definition video which makes it possible to emulate professional multiclip editing techniques for the case of single camera recordings. OpenKinoAI includes tools for…

Multimedia · Computer Science 2020-11-11 Rémi Ronfard , Rémi Colin de Verdière

Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow-based generation due to text-visual token imbalance and the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Jiabin Luo , Junhui Lin , Zeyu Zhang , Biao Wu , Meng Fang , Ling Chen , Hao Tang

Diffusion models have marked a significant milestone in the enhancement of image and video generation technologies. However, generating videos that precisely retain the shape and location of moving objects such as robots remains a…

Robotics · Computer Science 2024-07-04 Peng Wang , Zhihao Guo , Abdul Latheef Sait , Minh Huy Pham

When editing a video, a piece of attractive background music is indispensable. However, video background music generation tasks face several challenges, for example, the lack of suitable training datasets, and the difficulties in flexibly…

Computer Vision and Pattern Recognition · Computer Science 2024-05-21 Sizhe Li , Yiming Qin , Minghang Zheng , Xin Jin , Yang Liu

The flourishing of video generation technologies has endangered the credibility of real-world information and intensified the demand for AI-generated video detectors. Despite some progress, the lack of high-quality real-world datasets…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Weiliang Chen , Wenzhao Zheng , Yu Zheng , Lei Chen , Jie Zhou , Jiwen Lu , Yueqi Duan

Recent advances in diffusion models bring new vitality to visual content creation. However, current text-to-video generation models still face significant challenges such as high training costs, substantial data requirements, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Sicong Feng , Jielong Yang , Li Peng

Generating consistent long videos is a complex challenge: while diffusion-based generative models generate visually impressive short clips, extending them to longer durations often leads to memory bottlenecks and long-term inconsistency. In…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Wenqi Ouyang , Zeqi Xiao , Danni Yang , Yifan Zhou , Shuai Yang , Lei Yang , Jianlou Si , Xingang Pan

Recent work like GPT-3 has demonstrated excellent performance of Zero-Shot and Few-Shot learning on many natural language processing (NLP) tasks by scaling up model size, dataset size and the amount of computation. However, training a model…

Computation and Language · Computer Science 2021-10-13 Shaohua Wu , Xudong Zhao , Tong Yu , Rongguo Zhang , Chong Shen , Hongli Liu , Feng Li , Hong Zhu , Jiangang Luo , Liang Xu , Xuanwei Zhang

Generating higher-resolution human-centric scenes with details and controls remains a challenge for existing text-to-image diffusion models. This challenge stems from limited training image size, text encoder capacity (limited tokens), and…

Computer Vision and Pattern Recognition · Computer Science 2024-04-09 Gwanghyun Kim , Hayeon Kim , Hoigi Seo , Dong Un Kang , Se Young Chun