中文
相关论文

相关论文: VideoUFO: A Million-Scale User-Focused Dataset for…

200 篇论文

We introduce a novel diffusion-based video generation method, generating a video showing multiple events given multiple individual sentences from the user. Our method does not require a large-scale video dataset since our method uses a…

计算机视觉与模式识别 · 计算机科学 2024-07-17 Gyeongrok Oh , Jaehwan Jeong , Sieun Kim , Wonmin Byeon , Jinkyu Kim , Sungwoong Kim , Sangpil Kim

Automatically identifying harmful content in video is an important task with a wide range of applications. However, there is a lack of professionally labeled open datasets available. In this work VidHarm, an open dataset of 3589 video clips…

计算机视觉与模式识别 · 计算机科学 2022-09-07 Johan Edstedt , Amanda Berg , Michael Felsberg , Johan Karlsson , Francisca Benavente , Anette Novak , Gustav Grund Pihlgren

Recent advances in text-to-video (T2V) generation have achieved good visual quality, yet synthesizing videos that faithfully follow physical laws remains an open challenge. Existing methods mainly based on graphics or prompt extension…

计算机视觉与模式识别 · 计算机科学 2026-03-06 Yuanhao Cai , Kunpeng Li , Menglin Jia , Jialiang Wang , Junzhe Sun , Feng Liang , Weifeng Chen , Felix Juefei-Xu , Chu Wang , Ali Thabet , Xiaoliang Dai , Xuan Ju , Alan Yuille , Ji Hou

We propose a novel Unmanned Aerial Vehicles (UAV) assisted creative capture system that leverages diffusion models to interpret high-level natural language prompts and automatically generate optimal flight trajectories for cinematic video…

机器人学 · 计算机科学 2026-01-27 Valerii Serpiva , Artem Lykov , Jeffrin Sam , Aleksey Fedoseev , Dzmitry Tsetserukou

Recent advances in text-to-video (T2V) technology, as demonstrated by models such as Runway Gen-3, Pika, Sora, and Kling, have significantly broadened the applicability and popularity of the technology. This progress has created a growing…

计算机视觉与模式识别 · 计算机科学 2026-01-27 Zelu Qi , Ping Shi , Shuqi Wang , Chaoyang Zhang , Fei Zhao , Zefeng Ying , Da Pan , Xi Yang , Zheqi He , Teng Dai

Constructing supervised machine learning models for real-world video analysis require substantial labeled data, which is costly to acquire due to scarce domain expertise and laborious manual inspection. While data programming shows promise…

计算机视觉与模式识别 · 计算机科学 2023-11-02 Jianben He , Xingbo Wang , Kam Kwai Wong , Xijie Huang , Changjian Chen , Zixin Chen , Fengjie Wang , Min Zhu , Huamin Qu

Recent advances in video generation have shown remarkable progress in open-domain settings, yet medical video generation remains largely underexplored. Medical videos are critical for applications such as clinical training, education, and…

计算机视觉与模式识别 · 计算机科学 2025-07-09 Rongsheng Wang , Junying Chen , Ke Ji , Zhenyang Cai , Shunian Chen , Yunjin Yang , Benyou Wang

The recent advance in vision-language models is largely attributed to the abundance of image-text data. We aim to replicate this success for video-language models, but there simply is not enough human-curated video-text data available. We…

计算机视觉与模式识别 · 计算机科学 2024-04-17 Yue Zhao , Long Zhao , Xingyi Zhou , Jialin Wu , Chun-Te Chu , Hui Miao , Florian Schroff , Hartwig Adam , Ting Liu , Boqing Gong , Philipp Krähenbühl , Liangzhe Yuan

Text-to-video (T2V) generation has made tremendous progress in generating complicated scenes based on texts. However, human-object interaction (HOI) often cannot be precisely generated by current T2V models due to the lack of large-scale…

计算机视觉与模式识别 · 计算机科学 2025-04-01 Kun Liu , Qi Liu , Xinchen Liu , Jie Li , Yongdong Zhang , Jiebo Luo , Xiaodong He , Wu Liu

We propose a new two-stage pre-training framework for video-to-text generation tasks such as video captioning and video question answering: A generative encoder-decoder model is first jointly pre-trained on massive image-text data to learn…

计算机视觉与模式识别 · 计算机科学 2023-05-08 Xilun Chen , Lili Yu , Wenhan Xiong , Barlas Oğuz , Yashar Mehdad , Wen-tau Yih

We present a novel human annotated dataset for evaluating the ability for visual-language models to generate both short and long descriptions for real-world video clips, termed DeVAn (Dense Video Annotation). The dataset contains 8.5K…

计算机视觉与模式识别 · 计算机科学 2024-08-12 Tingkai Liu , Yunzhe Tao , Haogeng Liu , Qihang Fan , Ding Zhou , Huaibo Huang , Ran He , Hongxia Yang

Generating videos from text is a challenging task due to its high computational requirements for training and infinite possible answers for evaluation. Existing works typically experiment on simple or small datasets, where the…

计算机视觉与模式识别 · 计算机科学 2021-05-03 Chenfei Wu , Lun Huang , Qianxi Zhang , Binyang Li , Lei Ji , Fan Yang , Guillermo Sapiro , Nan Duan

Vision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections…

计算机视觉与模式识别 · 计算机科学 2023-10-10 Sihan Chen , Handong Li , Qunbo Wang , Zijia Zhao , Mingzhen Sun , Xinxin Zhu , Jing Liu

Real-world user-generated videos, especially on platforms like TikTok, often feature rich and intertwined audio visual content. However, existing video captioning benchmarks and models remain predominantly visual centric, overlooking the…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Peiran Wu , Yunze Liu , Zhengdong Zhu , Enmin Zhou , Junxiao Shen

While diffusion models have shown impressive performance in 2D image/video generation, diffusion-based Text-to-Multi-view-Video (T2MVid) generation remains underexplored. The new challenges posed by T2MVid generation lie in the lack of…

计算机视觉与模式识别 · 计算机科学 2024-06-14 Bing Li , Cheng Zheng , Wenxuan Zhu , Jinjie Mai , Biao Zhang , Peter Wonka , Bernard Ghanem

This paper presents \emph{ControlVideo} for text-driven video editing -- generating a video that aligns with a given text while preserving the structure of the source video. Building on a pre-trained text-to-image diffusion model,…

计算机视觉与模式识别 · 计算机科学 2023-11-29 Min Zhao , Rongzhen Wang , Fan Bao , Chongxuan Li , Jun Zhu

Video summarization aims to distill the most important information from a source video to produce either an abridged clip or a textual narrative. Traditionally, different methods have been proposed depending on whether the output is a video…

计算机视觉与模式识别 · 计算机科学 2024-04-24 Jingyang Lin , Hang Hua , Ming Chen , Yikang Li , Jenhao Hsiao , Chiuman Ho , Jiebo Luo

Recent methods for visual question answering rely on large-scale annotated datasets. Manual annotation of questions and answers for videos, however, is tedious, expensive and prevents scalability. In this work, we propose to avoid manual…

计算机视觉与模式识别 · 计算机科学 2022-05-12 Antoine Yang , Antoine Miech , Josef Sivic , Ivan Laptev , Cordelia Schmid

Text-driven diffusion models have unlocked unprecedented abilities in image generation, whereas their video counterpart still lags behind due to the excessive training cost of temporal modeling. Besides the training burden, the generated…

计算机视觉与模式识别 · 计算机科学 2023-05-23 Yabo Zhang , Yuxiang Wei , Dongsheng Jiang , Xiaopeng Zhang , Wangmeng Zuo , Qi Tian

Despite the number of currently available datasets on video question answering, there still remains a need for a dataset involving multi-step and non-factoid answers. Moreover, relying on video transcripts remains an under-explored topic.…

计算与语言 · 计算机科学 2020-06-02 Anthony Colas , Seokhwan Kim , Franck Dernoncourt , Siddhesh Gupte , Daisy Zhe Wang , Doo Soon Kim