English
Related papers

Related papers: RACCooN: A Versatile Instructional Video Editing F…

200 papers

Image-to-Video (I2V) generation aims to synthesize a video clip according to a given image and condition (e.g., text). The key challenge of this task lies in simultaneously generating natural motions while preserving the original appearance…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Jie Tian , Xiaoye Qu , Zhenyi Lu , Wei Wei , Sichen Liu , Yu Cheng

Video Generation is a relatively new and yet popular subject in machine learning due to its vast variety of potential applications and its numerous challenges. Current methods in Video Generation provide the user with little or no control…

Computer Vision and Pattern Recognition · Computer Science 2021-11-22 Bahman Rouhani , Mohammad Rahmati

While large-scale datasets have driven significant progress in Text-to-Video (T2V) generative models, these models remain highly sensitive to input prompts, demonstrating that prompt design is critical to generation quality. Current methods…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Zillur Rahman , Alex Sheng , Cristian Meo

Despite being (pre)trained on a massive amount of data, state-of-the-art video-language alignment models are not robust to semantically-plausible contrastive changes in the video captions. Our work addresses this by identifying a broad…

Computer Vision and Pattern Recognition · Computer Science 2023-11-20 Hritik Bansal , Yonatan Bitton , Idan Szpektor , Kai-Wei Chang , Aditya Grover

The rising popularity of immersive visual experiences has increased interest in stereoscopic 3D video generation. Despite significant advances in video synthesis, creating 3D videos remains challenging due to the relative scarcity of 3D…

Computer Vision and Pattern Recognition · Computer Science 2025-05-02 Michal Geyer , Omer Tov , Linyi Jin , Richard Tucker , Inbar Mosseri , Tali Dekel , Noah Snavely

Reference-to-video (R2V) generation is a controllable video synthesis paradigm that constrains the generation process using both text prompts and reference images, enabling applications such as personalized advertising and virtual try-on.…

Computer Vision and Pattern Recognition · Computer Science 2026-03-27 Lei Wang , YuXin Song , Ge Wu , Haocheng Feng , Hang Zhou , Jingdong Wang , Yaxing Wang , jian Yang

Recent great advances in video generation models have demonstrated their potential to produce high-quality videos, bringing challenges to effective evaluation. Unlike human evaluation, existing automated evaluation metrics lack highlevel…

Computer Vision and Pattern Recognition · Computer Science 2025-12-23 Zhun Mou , Bin Xia , Zhengchao Huang , Wenming Yang , Jiaya Jia

We introduce Vid-CamEdit, a novel framework for video camera trajectory editing, enabling the re-synthesis of monocular videos along user-defined camera paths. This task is challenging due to its ill-posed nature and the limited multi-view…

Computer Vision and Pattern Recognition · Computer Science 2025-06-17 Junyoung Seo , Jisang Han , Jaewoo Jung , Siyoon Jin , Joungbin Lee , Takuya Narihira , Kazumi Fukuda , Takashi Shibuya , Donghoon Ahn , Shoukang Hu , Seungryong Kim , Yuki Mitsufuji

Most methods for conditional video synthesis use a single modality as the condition. This comes with major limitations. For example, it is problematic for a model conditioned on an image to generate a specific motion trajectory desired by…

Computer Vision and Pattern Recognition · Computer Science 2022-03-08 Ligong Han , Jian Ren , Hsin-Ying Lee , Francesco Barbieri , Kyle Olszewski , Shervin Minaee , Dimitris Metaxas , Sergey Tulyakov

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale…

Computer Vision and Pattern Recognition · Computer Science 2025-04-01 Kun Liu , Qi Liu , Xinchen Liu , Jie Li , Yongdong Zhang , Jiebo Luo , Xiaodong He , Wu Liu

Video generation models have achieved remarkable progress in text-to-video tasks. These models are typically trained on text-video pairs with highly detailed and carefully crafted descriptions, while real-world user inputs during inference…

Computer Vision and Pattern Recognition · Computer Science 2025-09-03 Jiale Cheng , Ruiliang Lyu , Xiaotao Gu , Xiao Liu , Jiazheng Xu , Yida Lu , Jiayan Teng , Zhuoyi Yang , Yuxiao Dong , Jie Tang , Hongning Wang , Minlie Huang

Instructional video editing applies edits to an input video using only text prompts, enabling intuitive natural-language control. Despite rapid progress, most methods still require fixed-length inputs and substantial compute. Meanwhile,…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Mohammadreza Salehi , Mehdi Noroozi , Luca Morreale , Ruchika Chavhan , Malcolm Chadwick , Alberto Gil Ramos , Abhinav Mehrotra

Due to the rapid emergence of short videos and the requirement for content understanding and creation, the video captioning task has received increasing attention in recent years. In this paper, we convert traditional video captioning task…

Computer Vision and Pattern Recognition · Computer Science 2021-03-10 Ziqi Zhang , Zhongang Qi , Chunfeng Yuan , Ying Shan , Bing Li , Ying Deng , Weiming Hu

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

Computer Vision and Pattern Recognition · Computer Science 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee

Audiovisual video captioning aims to generate semantically rich descriptions with temporal alignment between visual and auditory events, thereby benefiting both video understanding and generation. In this paper, we present AVoCaDO, a…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Xinlong Chen , Yue Ding , Weihong Lin , Jingyun Hua , Linli Yao , Yang Shi , Bozhou Li , Yuanxing Zhang , Qiang Liu , Pengfei Wan , Liang Wang , Tieniu Tan

In this paper, we propose Text2Scene, a model that generates various forms of compositional scene representations from natural language descriptions. Unlike recent works, our method does NOT use Generative Adversarial Networks (GANs).…

Computer Vision and Pattern Recognition · Computer Science 2019-06-11 Fuwen Tan , Song Feng , Vicente Ordonez

Real-world videos naturally portray complex interactions among distinct physical objects, effectively forming dynamic compositions of visual elements. However, most current video generation models synthesize scenes holistically and…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Guofeng Zhang , Angtian Wang , Jacob Zhiyuan Fang , Liming Jiang , Haotian Yang , Alan Yuille , Chongyang Ma

A natural image usually conveys rich semantic content and can be viewed from different angles. Existing image description methods are largely restricted by small sets of biased visual paragraph annotations, and fail to cover rich underlying…

Computer Vision and Pattern Recognition · Computer Science 2017-03-27 Xiaodan Liang , Zhiting Hu , Hao Zhang , Chuang Gan , Eric P. Xing

The continuous development of foundational models for video generation is evolving into various applications, with subject-consistent video generation still in the exploratory stage. We refer to this as Subject-to-Video, which extracts…

Computer Vision and Pattern Recognition · Computer Science 2025-04-11 Lijie Liu , Tianxiang Ma , Bingchuan Li , Zhuowei Chen , Jiawei Liu , Gen Li , Siyu Zhou , Qian He , Xinglong Wu

We present Step-Video-T2V, a state-of-the-art text-to-video pre-trained model with 30B parameters and the ability to generate videos up to 204 frames in length. A deep compression Variational Autoencoder, Video-VAE, is designed for video…

Computer Vision and Pattern Recognition · Computer Science 2025-02-25 Guoqing Ma , Haoyang Huang , Kun Yan , Liangyu Chen , Nan Duan , Shengming Yin , Changyi Wan , Ranchen Ming , Xiaoniu Song , Xing Chen , Yu Zhou , Deshan Sun , Deyu Zhou , Jian Zhou , Kaijun Tan , Kang An , Mei Chen , Wei Ji , Qiling Wu , Wen Sun , Xin Han , Yanan Wei , Zheng Ge , Aojie Li , Bin Wang , Bizhu Huang , Bo Wang , Brian Li , Changxing Miao , Chen Xu , Chenfei Wu , Chenguang Yu , Dapeng Shi , Dingyuan Hu , Enle Liu , Gang Yu , Ge Yang , Guanzhe Huang , Gulin Yan , Haiyang Feng , Hao Nie , Haonan Jia , Hanpeng Hu , Hanqi Chen , Haolong Yan , Heng Wang , Hongcheng Guo , Huilin Xiong , Huixin Xiong , Jiahao Gong , Jianchang Wu , Jiaoren Wu , Jie Wu , Jie Yang , Jiashuai Liu , Jiashuo Li , Jingyang Zhang , Junjing Guo , Junzhe Lin , Kaixiang Li , Lei Liu , Lei Xia , Liang Zhao , Liguo Tan , Liwen Huang , Liying Shi , Ming Li , Mingliang Li , Muhua Cheng , Na Wang , Qiaohui Chen , Qinglin He , Qiuyan Liang , Quan Sun , Ran Sun , Rui Wang , Shaoliang Pang , Shiliang Yang , Sitong Liu , Siqi Liu , Shuli Gao , Tiancheng Cao , Tianyu Wang , Weipeng Ming , Wenqing He , Xu Zhao , Xuelin Zhang , Xianfang Zeng , Xiaojia Liu , Xuan Yang , Yaqi Dai , Yanbo Yu , Yang Li , Yineng Deng , Yingming Wang , Yilei Wang , Yuanwei Lu , Yu Chen , Yu Luo , Yuchu Luo , Yuhe Yin , Yuheng Feng , Yuxiang Yang , Zecheng Tang , Zekai Zhang , Zidong Yang , Binxing Jiao , Jiansheng Chen , Jing Li , Shuchang Zhou , Xiangyu Zhang , Xinhao Zhang , Yibo Zhu , Heung-Yeung Shum , Daxin Jiang