English
Related papers

Related papers: UniVG: Towards UNIfied-modal Video Generation

200 papers

Unified generation models aim to handle diverse tasks across modalities -- such as text generation, image generation, and vision-language reasoning -- within a single architecture and decoding paradigm. Autoregressive unified models suffer…

Machine Learning · Computer Science 2026-05-27 Qingyu Shi , Jinbin Bai , Zhuoran Zhao , Wenhao Chai , Kaidong Yu , Jianzong Wu , Yunhai Tong , Xiangtai Li , Xuelong Li , Shuicheng Yan

Despite diffusion models having shown powerful abilities to generate photorealistic images, generating videos that are realistic and diverse still remains in its infancy. One of the key reasons is that current methods intertwine spatial…

Computer Vision and Pattern Recognition · Computer Science 2023-12-08 Zhiwu Qing , Shiwei Zhang , Jiayu Wang , Xiang Wang , Yujie Wei , Yingya Zhang , Changxin Gao , Nong Sang

While image manipulation achieves tremendous breakthroughs (e.g., generating realistic faces) in recent years, video generation is much less explored and harder to control, which limits its applications in the real world. For instance,…

Computer Vision and Pattern Recognition · Computer Science 2019-08-08 Tsun-Hsuan Wang , Yen-Chi Cheng , Chieh Hubert Lin , Hwann-Tzong Chen , Min Sun

In recent years, significant progress has been made in the development of text-to-image generation models. However, these models still face limitations when it comes to achieving full controllability during the generation process. Often,…

Computer Vision and Pattern Recognition · Computer Science 2024-03-05 Salaheldin Mohamed

With the increasing popularity of autonomous driving based on the powerful and unified bird's-eye-view (BEV) representation, a demand for high-quality and large-scale multi-view video data with accurate annotation is urgently required.…

Computer Vision and Pattern Recognition · Computer Science 2023-10-13 Xiaofan Li , Yifu Zhang , Xiaoqing Ye

Recent advances in diffusion models bring new vitality to visual content creation. However, current text-to-video generation models still face significant challenges such as high training costs, substantial data requirements, and…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Sicong Feng , Jielong Yang , Li Peng

With the explosive popularity of AI-generated content (AIGC), video generation has recently received a lot of attention. Generating videos guided by text instructions poses significant challenges, such as modeling the complex relationship…

Computer Vision and Pattern Recognition · Computer Science 2024-04-25 Wenjing Wang , Huan Yang , Zixi Tuo , Huiguo He , Junchen Zhu , Jianlong Fu , Jiaying Liu

Long-term Video Question Answering (VideoQA) is a challenging vision-and-language bridging task focusing on semantic understanding of untrimmed long-term videos and diverse free-form questions, simultaneously emphasizing comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2024-10-15 Ting Yu , Kunhao Fu , Jian Zhang , Qingming Huang , Jun Yu

While text-to-3D and image-to-3D generation tasks have received considerable attention, one important but under-explored field between them is controllable text-to-3D generation, which we mainly focus on in this work. To address this task,…

Computer Vision and Pattern Recognition · Computer Science 2025-02-11 Zhiqi Li , Yiming Chen , Lingzhe Zhao , Peidong Liu

Generating multi-camera street-view videos is critical for augmenting autonomous driving datasets, addressing the urgent demand for extensive and varied data. Due to the limitations in diversity and challenges in handling lighting…

Computer Vision and Pattern Recognition · Computer Science 2024-08-09 Jiachen Lu , Ze Huang , Zeyu Yang , Jiahui Zhang , Li Zhang

Diffusion generative models have recently greatly improved the power of text-conditioned image generation. Existing image generation models mainly include text conditional diffusion model and cross-modal guided diffusion model, which are…

Computer Vision and Pattern Recognition · Computer Science 2022-11-04 Wei Li , Xue Xu , Xinyan Xiao , Jiachen Liu , Hu Yang , Guohao Li , Zhanpeng Wang , Zhifan Feng , Qiaoqiao She , Yajuan Lyu , Hua Wu

While state-of-the-art audio-video generation models like Veo3 and Sora2 demonstrate remarkable capabilities, their closed-source nature makes their architectures and training paradigms inaccessible. To bridge this gap in accessibility and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Hebeizi Li , Zihao Liang , Benyuan Sun , Zihao Yin , Xiao Sha , Chenliang Wang , Yi Yang

The field of text-to-3D content generation has made significant progress in generating realistic 3D objects, with existing methodologies like Score Distillation Sampling (SDS) offering promising guidance. However, these methods often…

Computer Vision and Pattern Recognition · Computer Science 2024-09-11 Phu Pham , Aradhya N. Mathur , Ojaswa Sharma , Aniket Bera

Multimodal large models have been recognized for their advantages in various performance and downstream tasks. The development of these models is crucial towards achieving general artificial intelligence in the future. In this paper, we…

Sound · Computer Science 2023-09-12 Sen Fang , Bowen Gao , Yangjian Wu , Teik Toe Teoh

Existing methods for vision-and-language learning typically require designing task-specific architectures and objectives for each task. For example, a multi-label answer classifier for visual question answering, a region scorer for…

Computation and Language · Computer Science 2021-05-25 Jaemin Cho , Jie Lei , Hao Tan , Mohit Bansal

Recent advancements in diffusion models have revolutionized video generation, enabling the creation of high-quality, temporally consistent videos. However, generating high frame-rate (FPS) videos remains a significant challenge due to…

Computer Vision and Pattern Recognition · Computer Science 2025-06-03 Geunmin Hwang , Hyun-kyu Ko , Younghyun Kim , Seungryong Lee , Eunbyung Park

Video grounding aims to locate the timestamps best matching the query description within an untrimmed video. Prevalent methods can be divided into moment-level and clip-level frameworks. Moment-level approaches directly predict the…

Computer Vision and Pattern Recognition · Computer Science 2023-03-15 Xing Cheng , Xiangyu Wu , Dong Shen , Hezheng Lin , Fan Yang

We introduce UniGen, a unified multimodal large language model (MLLM) capable of image understanding and generation. We study the full training pipeline of UniGen from a data-centric perspective, including multi-stage pre-training,…

Computer Vision and Pattern Recognition · Computer Science 2025-05-21 Rui Tian , Mingfei Gao , Mingze Xu , Jiaming Hu , Jiasen Lu , Zuxuan Wu , Yinfei Yang , Afshin Dehghan

Text-to-video generation is expensive, so only a few samples are typically produced per prompt. In this low-sample regime, maximizing the value of each batch requires high cross-video diversity. Recent methods improve diversity for image…

Computer Vision and Pattern Recognition · Computer Science 2026-02-18 Xinshuang Liu , Runfa Blark Li , Truong Nguyen

Video foundation models aim to integrate video understanding, generation, editing, and instruction following within a single framework, making them a central direction for next-generation multimodal systems. However, existing evaluation…

Computer Vision and Pattern Recognition · Computer Science 2026-03-09 Jianhui Wei , Xiaotian Zhang , Yichen Li , Yuan Wang , Yan Zhang , Ziyi Chen , Zhihang Tang , Wei Xu , Zuozhu Liu
‹ Prev 1 8 9 10 Next ›