English
Related papers

Related papers: GeneVA: A Dataset of Human Annotations for Generat…

200 papers

Emotion plays a pivotal role in video-based expression, but existing video generation systems predominantly focus on low-level visual metrics while neglecting affective dimensions. Although emotion analysis has made progress in the visual…

Computer Vision and Pattern Recognition · Computer Science 2025-11-17 Zongyang Qiu , Bingyuan Wang , Xingbei Chen , Yingqing He , Zeyu Wang

Recent advances in text-to-image generation have enabled the creation of high-quality images with diverse applications. However, accurately describing desired visual attributes can be challenging, especially for non-experts in art and…

Computer Vision and Pattern Recognition · Computer Science 2024-12-11 Tong Wu , Yinghao Xu , Ryan Po , Mengchen Zhang , Guandao Yang , Jiaqi Wang , Ziwei Liu , Dahua Lin , Gordon Wetzstein

Modern machine learning methods require significant amounts of labelled data, making the preparation process time-consuming and resource-intensive. In this paper, we propose to consider the process of prototyping a tool for annotating and…

Computer Vision and Pattern Recognition · Computer Science 2025-05-26 Nikita Ivanov , Mark Klimov , Dmitry Glukhikh , Tatiana Chernysheva , Igor Glukhikh

We consider the task of generating diverse and realistic videos guided by natural audio samples from a wide variety of semantic classes. For this task, the videos are required to be aligned both globally and temporally with the input audio:…

Machine Learning · Computer Science 2023-09-29 Guy Yariv , Itai Gat , Sagie Benaim , Lior Wolf , Idan Schwartz , Yossi Adi

This work introduces a new task, text-conditioned selective video-to-audio (V2A) generation, which produces only the user-intended sound from a multi-object video. This capability is especially crucial in multimedia production, where audio…

Computer Vision and Pattern Recognition · Computer Science 2026-03-30 Junwon Lee , Juhan Nam , Jiyoung Lee

Story visualization is an under-explored task that falls at the intersection of many important research directions in both computer vision and natural language processing. In this task, given a series of natural language captions which…

Computation and Language · Computer Science 2021-05-24 Adyasha Maharana , Darryl Hannan , Mohit Bansal

Conversational generative vision models (CGVMs) like Visual ChatGPT (Wu et al., 2023) have recently emerged from the synthesis of computer vision and natural language processing techniques. These models enable more natural and interactive…

Computer Vision and Pattern Recognition · Computer Science 2023-05-30 Narjes Nikzad Khasmakhi , Meysam Asgari-Chenaghlu , Nabiha Asghar , Philipp Schaer , Dietlind Zühlke

AI-driven video generation techniques have made significant progress in recent years. However, AI-generated videos (AGVs) involving human activities often exhibit substantial visual and semantic distortions, hindering the practical…

Computer Vision and Pattern Recognition · Computer Science 2025-07-24 Zhichao Zhang , Wei Sun , Xinyue Li , Yunhao Li , Qihang Ge , Jun Jia , Zicheng Zhang , Zhongpeng Ji , Fengyu Sun , Shangling Jui , Xiongkuo Min , Guangtao Zhai

Visual and auditory perception are two crucial ways humans experience the world. Text-to-video generation has made remarkable progress over the past year, but the absence of harmonious audio in generated video limits its broader…

Sound · Computer Science 2025-03-25 Yong Ren , Chenxing Li , Manjie Xu , Wei Liang , Yu Gu , Rilin Chen , Dong Yu

We explore spatiotemporal data augmentation using video foundation models to diversify both camera viewpoints and scene dynamics. Unlike existing approaches based on simple geometric transforms or appearance perturbations, our method…

Computer Vision and Pattern Recognition · Computer Science 2025-12-16 Jinfan Zhou , Lixin Luo , Sungmin Eum , Heesung Kwon , Jeong Joon Park

Recent advances in deep generative models have lead to remarkable progress in synthesizing high quality images. Following their successful application in image processing and representation learning, an important next step is to consider…

Computer Vision and Pattern Recognition · Computer Science 2019-03-28 Thomas Unterthiner , Sjoerd van Steenkiste , Karol Kurach , Raphael Marinier , Marcin Michalski , Sylvain Gelly

Text-to-video (T2V) synthesis has advanced rapidly, yet current evaluation metrics primarily capture visual quality and temporal consistency, offering limited insight into how synthetic videos perform in downstream tasks such as…

Computer Vision and Pattern Recognition · Computer Science 2025-07-04 Zecheng Zhao , Selena Song , Tong Chen , Zhi Chen , Shazia Sadiq , Yadan Luo

Video Annotation is a crucial process in computer science and social science alike. Many video annotation tools (VATs) offer a wide range of features for making annotation possible. We conducted an extensive survey of over 59 VATs and…

Human-Computer Interaction · Computer Science 2023-01-10 Snehesh Shrestha , William Sentosatio , Huiashu Peng , Cornelia Fermuller , Yiannis Aloimonos

The field of video generation has expanded significantly in recent years, with controllable and compositional video generation garnering considerable interest. Most methods rely on leveraging annotations such as text, objects' bounding…

Computer Vision and Pattern Recognition · Computer Science 2025-03-25 Aram Davtyan , Sepehr Sameni , Björn Ommer , Paolo Favaro

Recent advancements in visual generation technologies have markedly increased the scale and availability of video datasets, which are crucial for training effective video generation models. However, a significant lack of high-quality,…

Computer Vision and Pattern Recognition · Computer Science 2025-01-07 Hui Li , Mingwang Xu , Yun Zhan , Shan Mu , Jiaye Li , Kaihui Cheng , Yuxuan Chen , Tan Chen , Mao Ye , Jingdong Wang , Siyu Zhu

We propose Make-A-Video -- an approach for directly translating the tremendous recent progress in Text-to-Image (T2I) generation to Text-to-Video (T2V). Our intuition is simple: learn what the world looks like and how it is described from…

Computer Vision and Pattern Recognition · Computer Science 2022-09-30 Uriel Singer , Adam Polyak , Thomas Hayes , Xi Yin , Jie An , Songyang Zhang , Qiyuan Hu , Harry Yang , Oron Ashual , Oran Gafni , Devi Parikh , Sonal Gupta , Yaniv Taigman

With the tremendous growth of videos over the Internet, video thumbnails, providing video content previews, are becoming increasingly crucial to influencing users' online searching experiences. Conventional video thumbnails are generated…

Computer Vision and Pattern Recognition · Computer Science 2019-10-17 Yitian Yuan , Lin Ma , Wenwu Zhu

With the growing popularity of text-to-image generative models, there has been increasing focus on understanding their risks and biases. Recent work has found that state-of-the-art models struggle to depict everyday objects with the true…

Computer Vision and Pattern Recognition · Computer Science 2024-08-05 Reyhane Askari Hemmat , Melissa Hall , Alicia Sun , Candace Ross , Michal Drozdzal , Adriana Romero-Soriano

There has been an explosion of work in the vision & language community during the past few years from image captioning to video transcription, and answering questions about images. These tasks have focused on literal descriptions of the…

Computation and Language · Computer Science 2016-06-10 Nasrin Mostafazadeh , Ishan Misra , Jacob Devlin , Margaret Mitchell , Xiaodong He , Lucy Vanderwende

Recent advancements in text-to-video models such as Sora, Gen-3, MovieGen, and CogVideoX are pushing the boundaries of synthetic video generation, with adoption seen in fields like robotics, autonomous driving, and entertainment. As these…

Computer Vision and Pattern Recognition · Computer Science 2025-04-28 S P Sharan , Minkyu Choi , Sahil Shah , Harsh Goel , Mohammad Omama , Sandeep Chinchali
‹ Prev 1 3 4 5 6 7 10 Next ›