English
Related papers

Related papers: HunyuanVideo 1.5 Technical Report

200 papers

Video generation models have demonstrated remarkable performance, yet their broader adoption remains constrained by slow inference speeds and substantial computational costs, primarily due to the iterative nature of the denoising process.…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Xin Zhou , Dingkang Liang , Kaijin Chen , Tianrui Feng , Xiwu Chen , Hongkai Lin , Yikang Ding , Feiyang Tan , Hengshuang Zhao , Xiang Bai

We present ACE-Step v1.5, a highly efficient open-source music foundation model that brings commercial-grade generation to consumer hardware. On commonly used evaluation metrics, ACE-Step v1.5 achieves quality beyond most commercial music…

Sound · Computer Science 2026-02-09 Junmin Gong , Yulin Song , Wenxiao Zhao , Sen Wang , Shengyuan Xu , Jing Guo , Xuerui Yang

This paper introduces Goku, a state-of-the-art family of joint image-and-video generation models leveraging rectified flow Transformers to achieve industry-leading performance. We detail the foundational elements enabling high-quality…

High-resolution video generation has emerged as a crucial task in computer vision, with wide-ranging applications in entertainment, simulation, and data augmentation. However, generating temporally coherent and visually realistic videos…

Image and Video Processing · Electrical Eng. & Systems 2025-07-08 Abhinav Sagar

Recent advances in video generation demand increasingly efficient training recipes to mitigate escalating computational costs. In this report, we present ContentV, an 8B-parameter text-to-video model that achieves state-of-the-art…

Computer Vision and Pattern Recognition · Computer Science 2025-06-12 Wenfeng Lin , Renjie Chen , Boyuan Liu , Shiyue Yan , Ruoyu Feng , Jiangchuan Wei , Yichen Zhang , Yimeng Zhou , Chao Feng , Jiao Ran , Qi Wu , Zuotao Liu , Mingyu Guo

Generating videos predicting the future of a given sequence has been an area of active research in recent years. However, an essential problem remains unsolved: most of the methods require large computational cost and memory usage for…

Computer Vision and Pattern Recognition · Computer Science 2021-06-09 Naoya Fushishita , Antonio Tejero-de-Pablos , Yusuke Mukuta , Tatsuya Harada

Video generation requires modeling a vast spatiotemporal space, which demands significant computational resources and data usage. To reduce the complexity, the prevailing approaches employ a cascaded architecture to avoid direct training…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Yang Jin , Zhicheng Sun , Ningyuan Li , Kun Xu , Kun Xu , Hao Jiang , Nan Zhuang , Quzhe Huang , Yang Song , Yadong Mu , Zhouchen Lin

We present Omni-Video 2, a scalable and computationally efficient model that connects pretrained multimodal large-language models (MLLMs) with video diffusion models for unified video generation and editing. Our key idea is to exploit the…

Computer Vision and Pattern Recognition · Computer Science 2026-03-16 Hao Yang , Zhiyu Tan , Jia Gong , Luozheng Qin , Hesen Chen , Xiaomeng Yang , Yuqing Sun , Yuetan Lin , Mengping Yang , Hao Li

We introduce TurboDiffusion, a video generation acceleration framework that can speed up end-to-end diffusion generation by 100-200x while maintaining video quality. TurboDiffusion mainly relies on several components for acceleration: (1)…

Computer Vision and Pattern Recognition · Computer Science 2025-12-19 Jintao Zhang , Kaiwen Zheng , Kai Jiang , Haoxu Wang , Ion Stoica , Joseph E. Gonzalez , Jianfei Chen , Jun Zhu

In this report, we introduce InternVL 1.5, an open-source multimodal large language model (MLLM) to bridge the capability gap between open-source and proprietary commercial models in multimodal understanding. We introduce three simple…

We present Step-Video-TI2V, a state-of-the-art text-driven image-to-video generation model with 30B parameters, capable of generating videos up to 102 frames based on both text and image inputs. We build Step-Video-TI2V-Eval as a new…

High-resolution video generation, while crucial for digital media and film, is computationally bottlenecked by the quadratic complexity of diffusion models, making practical inference infeasible. To address this, we introduce HiStream, an…

Computer Vision and Pattern Recognition · Computer Science 2025-12-29 Haonan Qiu , Shikun Liu , Zijian Zhou , Zhaochong An , Weiming Ren , Zhiheng Liu , Jonas Schult , Sen He , Shoufa Chen , Yuren Cong , Tao Xiang , Ziwei Liu , Juan-Manuel Perez-Rua

Seedance 2.0 is a new native multi-modal audio-video generation model, officially released in China in early February 2026. Compared with its predecessors, Seedance 1.0 and 1.5 Pro, Seedance 2.0 adopts a unified, highly efficient, and…

Computer Vision and Pattern Recognition · Computer Science 2026-04-16 Team Seedance , De Chen , Liyang Chen , Xin Chen , Ying Chen , Zhuo Chen , Zhuowei Chen , Feng Cheng , Tianheng Cheng , Yufeng Cheng , Mojie Chi , Xuyan Chi , Jian Cong , Qinpeng Cui , Fei Ding , Qide Dong , Yujiao Du , Haojie Duanmu , Junliang Fan , Jiarui Fang , Jing Fang , Zetao Fang , Chengjian Feng , Yu Gao , Diandian Gu , Dong Guo , Hanzhong Guo , Qiushan Guo , Boyang Hao , Hongxiang Hao , Haoxun He , Jiaao He , Qian He , Tuyen Hoang , Heng Hu , Ruoqing Hu , Yuxiang Hu , Jiancheng Huang , Weilin Huang , Zhaoyang Huang , Zhongyi Huang , Jishuo Jin , Ming Jing , Ashley Kim , Shanshan Lao , Yichong Leng , Bingchuan Li , Gen Li , Haifeng Li , Huixia Li , Jiashi Li , Ming Li , Xiaojie Li , Xingxing Li , Yameng Li , Yiying Li , Yu Li , Yueyan Li , Chao Liang , Han Liang , Jianzhong Liang , Ying Liang , Wang Liao , J. H. Lien , Shanchuan Lin , Xi Lin , Feng Ling , Yue Ling , Fangfang Liu , Jiawei Liu , Jihao Liu , Jingtuo Liu , Shu Liu , Sichao Liu , Wei Liu , Xue Liu , Zuxi Liu , Ruijie Lu , Lecheng Lyu , Jingting Ma , Tianxiang Ma , Xiaonan Nie , Jingzhe Ning , Junjie Pan , Xitong Pan , Ronggui Peng , Xueqiong Qu , Yuxi Ren , Yuchen Shen , Guang Shi , Lei Shi , Yinglong Song , Fan Sun , Li Sun , Renfei Sun , Wenjing Tang , Boyang Tao , Zirui Tao , Dongliang Wang , Feng Wang , Hulin Wang , Ke Wang , Qingyi Wang , Rui Wang , Shuai Wang , Shulei Wang , Weichen Wang , Xuanda Wang , Yanhui Wang , Yue Wang , Yuping Wang , Yuxuan Wang , Zijie Wang , Ziyu Wang , Guoqiang Wei , Meng Wei , Di Wu , Guohong Wu , Hanjie Wu , Huachao Wu , Jian Wu , Jie Wu , Ruolan Wu , Shaojin Wu , Xiaohu Wu , Xinglong Wu , Yonghui Wu , Ruiqi Xia , Xin Xia , Xuefeng Xiao , Shuang Xu , Bangbang Yang , Jiaqi Yang , Runkai Yang , Tao Yang , Yihang Yang , Zhixian Yang , Ziyan Yang , Fulong Ye , Bingqian Yi , Xing Yin , Yongbin You , Linxiao Yuan , Weihong Zeng , Xuejiao Zeng , Yan Zeng , Siyu Zhai , Zhonghua Zhai , Bowen Zhang , Chenlin Zhang , Heng Zhang , Jun Zhang , Manlin Zhang , Peiyuan Zhang , Shuo Zhang , Xiaohe Zhang , Xiaoying Zhang , Xinyan Zhang , Xinyi Zhang , Yichi Zhang , Zixiang Zhang , Haiyu Zhao , Huating Zhao , Liming Zhao , Yian Zhao , Guangcong Zheng , Jianbin Zheng , Xiaozheng Zheng , Zerong Zheng , Kuan Zhu , Feilong Zuo

We present HY-Motion 1.0, a series of state-of-the-art, large-scale, motion generation models capable of generating 3D human motions from textual descriptions. HY-Motion 1.0 represents the first successful attempt to scale up Diffusion…

Multimodal Large Language Models (MLLMs) are undergoing rapid progress and represent the frontier of AI development. However, their training and inference efficiency have emerged as a core bottleneck in making MLLMs more accessible and…

Recently, video generation has achieved significant rapid development based on superior text-to-image generation techniques. In this work, we propose a high fidelity framework for image-to-video generation, named AtomoVideo. Based on…

Computer Vision and Pattern Recognition · Computer Science 2024-03-06 Litong Gong , Yiran Zhu , Weijie Li , Xiaoyang Kang , Biao Wang , Tiezheng Ge , Bo Zheng

We introduce SANA-Video, a small diffusion model that can efficiently generate videos up to 720x1280 resolution and minute-length duration. SANA-Video synthesizes high-resolution, high-quality and long videos with strong text-video…

Image-to-video generation, which aims to generate a video starting from a given reference image, has drawn great attention. Existing methods try to extend pre-trained text-guided image diffusion models to image-guided video generation…

Computer Vision and Pattern Recognition · Computer Science 2024-09-17 Cong Wang , Jiaxi Gu , Panwen Hu , Songcen Xu , Hang Xu , Xiaodan Liang

Unified video modeling that combines generation and understanding capabilities is increasingly important but faces two key challenges: maintaining semantic faithfulness during flow-based generation due to text-visual token imbalance and the…

Computer Vision and Pattern Recognition · Computer Science 2025-10-01 Jiabin Luo , Junhui Lin , Zeyu Zhang , Biao Wu , Meng Fang , Ling Chen , Hao Tang

Video generation has been advancing rapidly, and diffusion transformer (DiT) based models have demonstrated remark- able capabilities. However, their practical deployment is of- ten hindered by slow inference speeds and high memory con-…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Sijie Wang , Qiang Wang , Shaohuai Shi
‹ Prev 1 3 4 5 6 7 10 Next ›