English
Related papers

Related papers: Wan: Open and Advanced Large-Scale Video Generativ…

200 papers

Diffusion Transformer has demonstrated powerful capability and scalability in generating high-quality images and videos. Further pursuing the unification of generation and editing tasks has yielded significant progress in the domain of…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Zeyinzi Jiang , Zhen Han , Chaojie Mao , Jingfeng Zhang , Yulin Pan , Yu Liu

One of the most significant challenges in statistical signal processing and machine learning is how to obtain a generative model that can produce samples of large-scale data distribution, such as images and speeches. Generative Adversarial…

Computer Vision and Pattern Recognition · Computer Science 2020-05-28 Pegah Salehi , Abdolah Chalechale , Maryam Taghizadeh

Document image enhancement and binarization are commonly performed prior to document analysis and recognition tasks for improving the efficiency and accuracy of optical character recognition (OCR) systems. This is because directly…

Computer Vision and Pattern Recognition · Computer Science 2025-12-17 Rui-Yang Ju , KokSheik Wong , Yanlin Jin , Jen-Shiun Chiang

Numerous text-to-video (T2V) editing methods have emerged recently, but the lack of a standardized benchmark for fair evaluation has led to inconsistent claims and an inability to assess model sensitivity to hyperparameters. Fine-grained…

Computer Vision and Pattern Recognition · Computer Science 2025-07-23 Minghan Li , Chenxi Xie , Yichen Wu , Lei Zhang , Mengyu Wang

We present W.A.L.T, a transformer-based approach for photorealistic video generation via diffusion modeling. Our approach has two key design decisions. First, we use a causal encoder to jointly compress images and videos within a unified…

Computer Vision and Pattern Recognition · Computer Science 2023-12-12 Agrim Gupta , Lijun Yu , Kihyuk Sohn , Xiuye Gu , Meera Hahn , Li Fei-Fei , Irfan Essa , Lu Jiang , José Lezama

Latent variable generative models have emerged as powerful tools for generative tasks including image and video synthesis. These models are enabled by pretrained autoencoders that map high resolution data into a compressed lower dimensional…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Mohammed Suhail , Carlos Esteves , Leonid Sigal , Ameesh Makadia

Multimodal deep neural networks deployed in realistic environments must contend with runtime variations: changes in modality quality, overall input complexity, and available platform resources. Current networks struggle with such…

Machine Learning · Computer Science 2026-05-04 Jason Wu , Shir-Kang Scott Jin , Yuyang Yuan , Maggie Wigness , Lance M. Kaplan , Hang Qiu , Mani Srivastava

We introduce Virtual Width Networks (VWN), a framework that delivers the benefits of wider representations without incurring the quadratic cost of increasing the hidden size. VWN decouples representational width from backbone width,…

Machine Learning · Computer Science 2025-11-18 Seed , Baisheng Li , Banggu Wu , Bole Ma , Bowen Xiao , Chaoyi Zhang , Cheng Li , Chengyi Wang , Chengyin Xu , Chi Zhang , Chong Hu , Daoguang Zan , Defa Zhu , Dongyu Xu , Du Li , Faming Wu , Fan Xia , Ge Zhang , Guang Shi , Haobin Chen , Hongyu Zhu , Hongzhi Huang , Huan Zhou , Huanzhang Dou , Jianhui Duan , Jianqiao Lu , Jianyu Jiang , Jiayi Xu , Jiecao Chen , Jin Chen , Jin Ma , Jing Su , Jingji Chen , Jun Wang , Jun Yuan , Juncai Liu , Jundong Zhou , Kai Hua , Kai Shen , Kai Xiang , Kaiyuan Chen , Kang Liu , Ke Shen , Liang Xiang , Lin Yan , Lishu Luo , Mengyao Zhang , Ming Ding , Mofan Zhang , Nianning Liang , Peng Li , Penghao Huang , Pengpeng Mu , Qi Huang , Qianli Ma , Qiyang Min , Qiying Yu , Renming Pang , Ru Zhang , Shen Yan , Shen Yan , Shixiong Zhao , Shuaishuai Cao , Shuang Wu , Siyan Chen , Siyu Li , Siyuan Qiao , Tao Sun , Tian Xin , Tiantian Fan , Ting Huang , Ting-Han Fan , Wei Jia , Wenqiang Zhang , Wenxuan Liu , Xiangzhong Wu , Xiaochen Zuo , Xiaoying Jia , Ximing Yang , Xin Liu , Xin Yu , Xingyan Bin , Xintong Hao , Xiongcai Luo , Xujing Li , Xun Zhou , Yanghua Peng , Yangrui Chen , Yi Lin , Yichong Leng , Yinghao Li , Yingshuan Song , Yiyuan Ma , Yong Shan , Yongan Xiang , Yonghui Wu , Yongtao Zhang , Yongzhen Yao , Yu Bao , Yuehang Yang , Yufeng Yuan , Yunshui Li , Yuqiao Xian , Yutao Zeng , Yuxuan Wang , Zehua Hong , Zehua Wang , Zengzhi Wang , Zeyu Yang , Zhengqiang Yin , Zhenyi Lu , Zhexi Zhang , Zhi Chen , Zhi Zhang , Zhiqi Lin , Zihao Huang , Zilin Xu , Ziyun Wei , Zuo Wang

Inspired by generative paradigms in image and video, 3D shape generation has made notable progress, enabling the rapid synthesis of high-fidelity 3D assets from a single image. However, current methods still face challenges, including the…

Computer Vision and Pattern Recognition · Computer Science 2025-11-26 Yangguang Li , Xianglong He , Zi-Xin Zou , Zexiang Liu , Wanli Ouyang , Ding Liang , Yan-Pei Cao

We present VideoGPT: a conceptually simple architecture for scaling likelihood based generative modeling to natural videos. VideoGPT uses VQ-VAE that learns downsampled discrete latent representations of a raw video by employing 3D…

Computer Vision and Pattern Recognition · Computer Science 2021-09-16 Wilson Yan , Yunzhi Zhang , Pieter Abbeel , Aravind Srinivas

While generative artificial intelligence has advanced significantly across text, image, audio, and video domains, 3D generation remains comparatively underdeveloped due to fundamental challenges such as data scarcity, algorithmic…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Weiyu Li , Xuanyang Zhang , Zheng Sun , Di Qi , Hao Li , Wei Cheng , Weiwei Cai , Shihao Wu , Jiarui Liu , Zihao Wang , Xiao Chen , Feipeng Tian , Jianxiong Pan , Zeming Li , Gang Yu , Xiangyu Zhang , Daxin Jiang , Ping Tan

This work introduces World-GAN, the first method to perform data-driven Procedural Content Generation via Machine Learning in Minecraft from a single example. Based on a 3D Generative Adversarial Network (GAN) architecture, we are able to…

Machine Learning · Computer Science 2021-06-21 Maren Awiszus , Frederik Schubert , Bodo Rosenhahn

The vision community is witnessing a modeling shift from CNNs to Transformers, where pure Transformer architectures have attained top accuracy on the major video recognition benchmarks. These video models are all built on Transformer layers…

Computer Vision and Pattern Recognition · Computer Science 2021-06-25 Ze Liu , Jia Ning , Yue Cao , Yixuan Wei , Zheng Zhang , Stephen Lin , Han Hu

As world models gain momentum in Embodied AI, an increasing number of works explore using video foundation models as predictive world models for downstream embodied tasks like 3D prediction or interactive generation. However, before…

This work aims to learn a high-quality text-to-video (T2V) generative model by leveraging a pre-trained text-to-image (T2I) model as a basis. It is a highly desirable yet challenging task to simultaneously a) accomplish the synthesis of…

Generative Adversarial Networks (GANs) are a powerful class of generative models. Despite their successes, the most appropriate choice of a GAN network architecture is still not well understood. GAN models for image synthesis have adopted a…

Machine Learning · Computer Science 2019-05-28 Sukarna Barua , Sarah Monazam Erfani , James Bailey

We introduce FewGAN, a generative model for generating novel, high-quality and diverse images whose patch distribution lies in the joint patch distribution of a small number of N>1 training samples. The method is, in essence, a hierarchical…

Computer Vision and Pattern Recognition · Computer Science 2022-07-25 Lior Ben-Moshe , Sagie Benaim , Lior Wolf

Diffusion models have recently achieved remarkable success in generative tasks (e.g., image and video generation), and the demand for high-quality content (e.g., 2K/4K videos) is rapidly increasing across various domains. However,…

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang