中文
相关论文

相关论文: VideoUFO: A Million-Scale User-Focused Dataset for…

200 篇论文

In this paper, we present VideoGen, a text-to-video generation approach, which can generate a high-definition video with high frame fidelity and strong temporal consistency using reference-guided latent diffusion. We leverage an…

计算机视觉与模式识别 · 计算机科学 2023-09-08 Xin Li , Wenqing Chu , Ye Wu , Weihang Yuan , Fanglong Liu , Qi Zhang , Fu Li , Haocheng Feng , Errui Ding , Jingdong Wang

As short-form funny videos on social networks are gaining popularity, it becomes demanding for AI models to understand them for better communication with humans. Unfortunately, previous video humor datasets target specific domains, such as…

计算与语言 · 计算机科学 2024-04-02 Dayoon Ko , Sangho Lee , Gunhee Kim

Video question answering (VideoQA) aims to answer natural language questions according to the given videos. Although existing models perform well in the factoid VideoQA task, they still face challenges in deep video understanding (DVU)…

计算机视觉与模式识别 · 计算机科学 2024-12-24 Zhengqian Wu , Ruizhe Li , Zijun Xu , Zhongyuan Wang , Chunxia Xiao , Chao Liang

Music shapes the tone of videos, yet creators often struggle to find soundtracks that match their video's mood and narrative. Recent text-to-music models let creators generate music from text prompts, but our formative study (N=8) shows…

人机交互 · 计算机科学 2026-02-05 Mina Huh , C. Ailie Fraser , Dingzeyu Li , Mira Dontcheva , Bryan Wang

The next frontier for video generation lies in developing models capable of zero-shot reasoning, where understanding real-world scientific laws is crucial for accurate physical outcome modeling under diverse conditions. However, existing…

计算机视觉与模式识别 · 计算机科学 2025-12-03 Lanxiang Hu , Abhilash Shankarampeta , Yixin Huang , Zilin Dai , Haoyang Yu , Yujie Zhao , Haoqiang Kang , Daniel Zhao , Tajana Rosing , Hao Zhang

Recent advances in video generation models demonstrate their potential as world simulators, but they often struggle with videos deviating from physical laws, a key concern overlooked by most text-to-video benchmarks. We introduce a…

计算机视觉与模式识别 · 计算机科学 2025-03-06 Yongfan Chen , Xiuwen Zhu , Tianyu Li

Text-to-video generation has demonstrated promising progress with the advent of diffusion models, yet existing approaches are limited by dataset quality and computational resources. To address these limitations, this paper presents a…

计算机视觉与模式识别 · 计算机科学 2025-03-10 Zhiyu Tan , Junyan Wang , Hao Yang , Luozheng Qin , Hesen Chen , Qiang Zhou , Hao Li

Educational videos have become increasingly relevant in today's learning environments. While prior research in laboratory studies has provided valuable insights, analyzing real-world interaction data can enhance our understanding of…

人机交互 · 计算机科学 2025-05-06 Wolfgang Gritz , Hewi Salih , Anett Hoppe , Ralph Ewerth

Text-to-video generation is an emerging field in generative AI, enabling the creation of realistic, semantically accurate videos from text prompts. While current models achieve impressive visual quality and alignment with input text, they…

计算机视觉与模式识别 · 计算机科学 2025-09-23 Luca Zanchetta , Lorenzo Papa , Luca Maiano , Irene Amerini

Facial images have extensive practical applications. Although the current large-scale text-image diffusion models exhibit strong generation capabilities, it is challenging to generate the desired facial images using only text prompt. Image…

计算机视觉与模式识别 · 计算机科学 2025-01-07 Dawei Dai , Mingming Jia , Yinxiu Zhou , Hang Xing , Chenghang Li

Despite the impressive progress of multimodal generative models, video-to-audio generation still suffers from limited performance and limits the flexibility to prioritize sound synthesis for specific objects within the scene. Conversely,…

计算机视觉与模式识别 · 计算机科学 2024-12-30 Yujin Jeong , Yunji Kim , Sanghyuk Chun , Jiyoung Lee

Text-to-video generation has evolved rapidly in recent years, delivering remarkable results. Training typically relies on video-caption paired data, which plays a crucial role in enhancing generation performance. However, current video…

计算机视觉与模式识别 · 计算机科学 2024-12-13 Tiehan Fan , Kepan Nan , Rui Xie , Penghao Zhou , Zhenheng Yang , Chaoyou Fu , Xiang Li , Jian Yang , Ying Tai

Many recent advancements in Computer Vision are attributed to large datasets. Open-source software packages for Machine Learning and inexpensive commodity hardware have reduced the barrier of entry for exploring novel approaches at scale.…

计算机视觉与模式识别 · 计算机科学 2016-09-29 Sami Abu-El-Haija , Nisarg Kothari , Joonseok Lee , Paul Natsev , George Toderici , Balakrishnan Varadarajan , Sudheendra Vijayanarasimhan

Vision-Language Models pre-trained on large-scale image-text datasets have shown superior performance in downstream tasks such as image retrieval. Most of the images for pre-training are presented in the form of open domain common-sense…

计算机视觉与模式识别 · 计算机科学 2024-01-26 Xiangshuo Qiao , Xianxin Li , Xiaozhe Qu , Jie Zhang , Yang Liu , Yu Luo , Cihang Jin , Jin Ma

Significant advancements in video diffusion models have brought substantial progress to the field of text-to-video (T2V) synthesis. However, existing T2V synthesis model struggle to accurately generate complex motion dynamics, leading to a…

计算机视觉与模式识别 · 计算机科学 2024-06-06 Haoran Cheng , Liang Peng , Linxuan Xia , Yuepeng Hu , Hengjia Li , Qinglin Lu , Xiaofei He , Boxi Wu

Advancements in multimodal learning, particularly in video understanding and generation, require high-quality video-text datasets for improved model performance. Vript addresses this issue with a meticulously annotated corpus of 12K…

计算机视觉与模式识别 · 计算机科学 2024-10-28 Dongjie Yang , Suyuan Huang , Chengqiang Lu , Xiaodong Han , Haoxin Zhang , Yan Gao , Yao Hu , Hai Zhao

Despite being (pre)trained on a massive amount of data, state-of-the-art video-language alignment models are not robust to semantically-plausible contrastive changes in the video captions. Our work addresses this by identifying a broad…

计算机视觉与模式识别 · 计算机科学 2023-11-20 Hritik Bansal , Yonatan Bitton , Idan Szpektor , Kai-Wei Chang , Aditya Grover

Learning long-term spatial-temporal features are critical for many video analysis tasks. However, existing video segmentation methods predominantly rely on static image segmentation techniques, and methods capturing temporal dependency for…

计算机视觉与模式识别 · 计算机科学 2018-09-11 Ning Xu , Linjie Yang , Yuchen Fan , Dingcheng Yue , Yuchen Liang , Jianchao Yang , Thomas Huang

Long video generation remains a challenging and compelling topic in computer vision. Diffusion based models, among the various approaches to video generation, have achieved state of the art quality with their iterative denoising procedures.…

计算机视觉与模式识别 · 计算机科学 2025-03-14 Siyang Zhang , Harry Yang , Ser-Nam Lim

With the recent popularity of animated GIFs on social media, there is need for ways to index them with rich metadata. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K animated GIFs…

计算机视觉与模式识别 · 计算机科学 2016-04-13 Yuncheng Li , Yale Song , Liangliang Cao , Joel Tetreault , Larry Goldberg , Alejandro Jaimes , Jiebo Luo