English
Related papers

Related papers: Wan: Open and Advanced Large-Scale Video Generativ…

200 papers

In recent years, multimodal large models have continued to improve on general benchmarks. However, in real-world content moderation and adversarial settings, mainstream models still suffer from degraded generalization and catastrophic…

Artificial Intelligence · Computer Science 2026-04-01 Zhiqian Zhang , Xu Zhao , Xiaoqing Xu , Guangdong Liang , Weijia Wang , Xiaolei Lv , Bo Li , Jun Gao

Generative classifiers, which leverage conditional generative models for classification, have recently demonstrated desirable properties such as robustness to distribution shifts. However, recent progress in this area has been largely…

Machine Learning · Computer Science 2026-03-24 Yi-Chung Chen , David I. Inouye , Jing Gao

Recent progress in diffusion models has greatly enhanced video generation quality, yet these models still require fine-tuning to improve specific dimensions like instance preservation, motion rationality, composition, and physical…

Computer Vision and Pattern Recognition · Computer Science 2025-06-13 Xiaoyi Bao , Jindi Lv , Xiaofeng Wang , Zheng Zhu , Xinze Chen , YuKun Zhou , Jiancheng Lv , Xingang Wang , Guan Huang

Contemporary models for generating images show remarkable quality and versatility. Swayed by these advantages, the research community repurposes them to generate videos. Since video content is highly redundant, we argue that naively…

Computer Vision and Pattern Recognition · Computer Science 2024-02-23 Willi Menapace , Aliaksandr Siarohin , Ivan Skorokhodov , Ekaterina Deyneka , Tsai-Shien Chen , Anil Kag , Yuwei Fang , Aleksei Stoliar , Elisa Ricci , Jian Ren , Sergey Tulyakov

World models have become crucial for autonomous driving, as they learn how scenarios evolve over time to address the long-tail challenges of the real world. However, current approaches relegate world models to limited roles: they operate…

Computer Vision and Pattern Recognition · Computer Science 2026-04-20 Tianze Xia , Yongkang Li , Lijun Zhou , Jingfeng Yao , Kaixin Xiong , Haiyang Sun , Bing Wang , Kun Ma , Guang Chen , Hangjun Ye , Wenyu Liu , Xinggang Wang

Large Language Models (LLMs) face a significant bottleneck during autoregressive inference due to the massive memory footprint of the Key-Value (KV) cache. Existing compression techniques like token eviction, quantization, or other low-rank…

Machine Learning · Computer Science 2025-11-25 Santhosh G S , Saurav Prakash , Balaraman Ravindran

Recent diffusion models enable high-quality video generation, but suffer from slow runtimes. The large transformer-based backbones used in these models are bottlenecked by spatiotemporal attention. In this paper, we identify that a…

Computer Vision and Pattern Recognition · Computer Science 2026-03-06 Shai Yehezkel , Shahar Yadin , Noam Elata , Yaron Ostrovsky-Berman , Bahjat Kawar

Efficient image tokenization with high compression ratios remains a critical challenge for training generative models. We present SoftVQ-VAE, a continuous image tokenizer that leverages soft categorical posteriors to aggregate multiple…

Computer Vision and Pattern Recognition · Computer Science 2025-03-18 Hao Chen , Ze Wang , Xiang Li , Ximeng Sun , Fangyi Chen , Jiang Liu , Jindong Wang , Bhiksha Raj , Zicheng Liu , Emad Barsoum

Video-based world models offer a powerful paradigm for embodied simulation and planning, yet state-of-the-art models often generate physically implausible manipulations - such as object penetration and anti-gravity motion - due to training…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Yuzhi Chen , Ronghan Chen , Dongjie Huo , Yandan Yang , Dekang Qi , Haoyun Liu , Tong Lin , Shuang Zeng , Junjin Xiao , Xinyuan Chang , Feng Xiong , Xing Wei , Zhiheng Ma , Mu Xu

Arbitrary resolution image generation provides a consistent visual experience across devices, having extensive applications for producers and consumers. Current diffusion models increase computational demand quadratically with resolution,…

Computer Vision and Pattern Recognition · Computer Science 2025-09-15 Tao Han , Wanghan Xu , Junchao Gong , Xiaoyu Yue , Song Guo , Luping Zhou , Lei Bai

Over the past years, deep generative models have achieved a new level of performance. Generated data has become difficult, if not impossible, to be distinguished from real data. While there are plenty of use cases that benefit from this…

Cryptography and Security · Computer Science 2022-03-21 Ning Yu , Vladislav Skripniuk , Dingfan Chen , Larry Davis , Mario Fritz

The usage of deep generative models for image compression has led to impressive performance gains over classical codecs while neural video compression is still in its infancy. Here, we propose an end-to-end, deep generative modeling…

Computer Vision and Pattern Recognition · Computer Science 2019-11-05 Jun Han , Salvator Lombardo , Christopher Schroers , Stephan Mandt

We introduce Sana, a text-to-image framework that can efficiently generate images up to 4096$\times$4096 resolution. Sana can synthesize high-resolution, high-quality images with strong text-image alignment at a remarkably fast speed,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Enze Xie , Junsong Chen , Junyu Chen , Han Cai , Haotian Tang , Yujun Lin , Zhekai Zhang , Muyang Li , Ligeng Zhu , Yao Lu , Song Han

We introduce a new encoder-decoder GAN model, FutureGAN, that predicts future frames of a video sequence conditioned on a sequence of past frames. During training, the networks solely receive the raw pixel values as an input, without…

Computer Vision and Pattern Recognition · Computer Science 2018-11-27 Sandra Aigner , Marco Körner

Recent success in deep learning has generated immense interest among practitioners and students, inspiring many to learn about this new technology. While visual and interactive approaches have been successfully developed to help people more…

Human-Computer Interaction · Computer Science 2018-09-06 Minsuk Kahng , Nikhil Thorat , Duen Horng Chau , Fernanda Viégas , Martin Wattenberg

Recent developments in Video Diffusion Models (VDMs) have demonstrated remarkable capability to generate high-quality video content. Nonetheless, the potential of VDMs for creating transparent videos remains largely uncharted. In this…

Graphics · Computer Science 2025-03-04 Menghao Li , Zhenghao Zhang , Junchao Liao , Long Qin , Weizhi Wang

Diffusion-based video generation techniques have significantly improved zero-shot talking-head avatar generation, enhancing the naturalness of both head motion and facial expressions. However, existing methods suffer from poor…

Graphics · Computer Science 2025-04-24 Lingzhou Mu , Baiji Liu , Ruonan Zhang , Guiming Mo , Jiawei Jin , Kai Zhang , Haozhi Huang

Generative models hold promise for revolutionizing medical education, robot-assisted surgery, and data augmentation for medical AI development. Diffusion models can now generate realistic images from text prompts, while recent advancements…

Computer Vision and Pattern Recognition · Computer Science 2024-07-17 Weixiang Sun , Xiaocao You , Ruizhe Zheng , Zhengqing Yuan , Xiang Li , Lifang He , Quanzheng Li , Lichao Sun

Endoscopic videos from multicentres often have different imaging conditions, e.g., color and illumination, which make the models trained on one domain usually fail to generalize well to another. Domain adaptation is one of the potential…

Computer Vision and Pattern Recognition · Computer Science 2020-04-20 Jiawei Chen , Yuexiang Li , Kai Ma , Yefeng Zheng

Generating flexible-view 3D scenes, including 360{\deg} rotation and zooming, from single images is challenging due to a lack of 3D data. To this end, we introduce FlexWorld, a novel framework consisting of two key components: (1) a strong…

Computer Vision and Pattern Recognition · Computer Science 2025-03-20 Luxi Chen , Zihan Zhou , Min Zhao , Yikai Wang , Ge Zhang , Wenhao Huang , Hao Sun , Ji-Rong Wen , Chongxuan Li
‹ Prev 1 8 9 10 Next ›