English
Related papers

Related papers: OpenVid-1M: A Large-Scale High-Quality Dataset for…

200 papers

Video generation models are rapidly advancing, but can still struggle with complex video outputs that require significant semantic branching or repeated high-level reasoning about what should happen next. In this paper, we introduce a new…

The generation of humanoid animation from text prompts can profoundly impact animation production and AR/VR experiences. However, existing methods only generate body motion data, excluding facial expressions and hand movements. This…

Computer Vision and Pattern Recognition · Computer Science 2024-09-23 Mingdian Liu , Yilin Liu , Gurunandan Krishnan , Karl S Bayer , Bing Zhou

In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Yang Jin , Zhicheng Sun , Kun Xu , Kun Xu , Liwei Chen , Hao Jiang , Quzhe Huang , Chengru Song , Yuliang Liu , Di Zhang , Yang Song , Kun Gai , Yadong Mu

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more…

High-resolution image-to-video (I2V) generation aims to synthesize realistic temporal dynamics while preserving fine-grained appearance details of the input image. At 2K resolution, it becomes extremely challenging, and existing solutions…

Computer Vision and Pattern Recognition · Computer Science 2026-05-12 YaoYang Liu , Yuechen Zhang , Wenbo Li , Yufei Zhao , Rui Liu , Long Chen

Generating multi-view images based on text or single-image prompts is a critical capability for the creation of 3D content. Two fundamental questions on this topic are what data we use for training and how to ensure multi-view consistency.…

Computer Vision and Pattern Recognition · Computer Science 2024-03-19 Qi Zuo , Xiaodong Gu , Lingteng Qiu , Yuan Dong , Zhengyi Zhao , Weihao Yuan , Rui Peng , Siyu Zhu , Zilong Dong , Liefeng Bo , Qixing Huang

Training large text-to-image models requires high-quality, curated datasets with diverse content and detailed captions. Yet the cost and complexity of collecting, filtering, deduplicating, and re-captioning such corpora at scale hinders…

Computer Vision and Pattern Recognition · Computer Science 2026-05-21 Benjamin Aubin , Gonzalo Iñaki Quintana , Onur Tasar , Sanjeev Sreetharan , Urszula Czerwinska , Damien Henry , Clément Chadebec

Existing multi-view image generation methods often make invasive modifications to pre-trained text-to-image (T2I) models and require full fine-tuning, leading to (1) high computational costs, especially with large base models and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-06 Zehuan Huang , Yuan-Chen Guo , Haoran Wang , Ran Yi , Lizhuang Ma , Yan-Pei Cao , Lu Sheng

Automated data visualization plays a crucial role in simplifying data interpretation, enhancing decision-making, and improving efficiency. While large language models (LLMs) have shown promise in generating visualizations from natural…

Computation and Language · Computer Science 2025-07-29 Mizanur Rahman , Md Tahmid Rahman Laskar , Shafiq Joty , Enamul Hoque

We present Waver, a high-performance foundation model for unified image and video generation. Waver can directly generate videos with durations ranging from 5 to 10 seconds at a native resolution of 720p, which are subsequently upscaled to…

Computer Vision and Pattern Recognition · Computer Science 2025-08-27 Yifu Zhang , Hao Yang , Yuqi Zhang , Yifei Hu , Fengda Zhu , Chuang Lin , Xiaofeng Mei , Yi Jiang , Bingyue Peng , Zehuan Yuan

Video generation is rapidly evolving towards unified audio-video generation. In this paper, we present ALIVE, a generation model that adapts a pretrained Text-to-Video (T2V) model to Sora-style audio-video generation and animation. In…

Computer Vision and Pattern Recognition · Computer Science 2026-02-11 Ying Guo , Qijun Gan , Yifu Zhang , Jinlai Liu , Yifei Hu , Pan Xie , Dongjun Qian , Yu Zhang , Ruiqi Li , Yuqi Zhang , Ruibiao Lu , Xiaofeng Mei , Bo Han , Xiang Yin , Bingyue Peng , Zehuan Yuan

While generative video models have achieved remarkable visual fidelity, their capacity to internalize and reason over implicit world rules remains a critical yet under-explored frontier. To bridge this gap, we present RISE-Video, a…

Computer Vision and Pattern Recognition · Computer Science 2026-02-06 Mingxin Liu , Shuran Ma , Shibei Meng , Xiangyu Zhao , Zicheng Zhang , Shaofeng Zhang , Zhihang Zhong , Peixian Chen , Haoyu Cao , Xing Sun , Haodong Duan , Xue Yang

In order to better simulate the real human conversation process, models need to generate dialogue utterances based on not only preceding textual contexts but also visual contexts. However, with the development of multi-modal dialogue…

Computation and Language · Computer Science 2021-09-29 Shuhe Wang , Yuxian Meng , Xiaoya Li , Xiaofei Sun , Rongbin Ouyang , Jiwei Li

Text to video generation has emerged as a critical frontier in generative artificial intelligence, yet existing approaches struggle with maintaining temporal consistency, compositional understanding, and fine grained control over visual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-04 Piyushkumar Patel

Text-to-3D generation approaches have advanced significantly by leveraging pretrained 2D diffusion priors, producing high-quality and 3D-consistent outputs. However, they often fail to produce out-of-domain (OOD) or rare concepts, yielding…

Computer Vision and Pattern Recognition · Computer Science 2025-08-25 Yosef Dayani , Omer Benishu , Sagie Benaim

Recent successful video generation systems that predict and create realistic automotive driving scenes from short video inputs assign tokenization, future state prediction (world model), and video decoding to dedicated models. These…

Computer Vision and Pattern Recognition · Computer Science 2025-09-22 Björn Möller , Zhengyang Li , Malte Stelzer , Thomas Graave , Fabian Bettels , Muaaz Ataya , Tim Fingscheidt

In text-to-video (T2V) generation, significant attention has been directed toward its development, yet unifying discrete and continuous grounding conditions in T2V generation remains under-explored. This paper proposes a Grounded…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Huanzhang Dou , Ruixiang Li , Wei Su , Xi Li

Over the past few years, Text-to-Image (T2I) generation approaches based on diffusion models have gained significant attention. However, vanilla diffusion models often suffer from spelling inaccuracies in the text displayed within the…

Computer Vision and Pattern Recognition · Computer Science 2024-10-30 Sanyam Lakhanpal , Shivang Chopra , Vinija Jain , Aman Chadha , Man Luo

Large-scale text-to-video (T2V) diffusion models have great progress in recent years in terms of visual quality, motion and temporal consistency. However, the generation process is still a black box, where all attributes (e.g., appearance,…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Jiwen Yu , Xiaodong Cun , Chenyang Qi , Yong Zhang , Xintao Wang , Ying Shan , Jian Zhang

Event-based cameras offer unique advantages such as high temporal resolution, high dynamic range, and low power consumption. However, the massive storage requirements and I/O burdens of existing synthetic data generation pipelines and the…

Computer Vision and Pattern Recognition · Computer Science 2025-05-23 Hanyue Lou , Jinxiu Liang , Minggui Teng , Yi Wang , Boxin Shi