English
Related papers

Related papers: VideoUFO: A Million-Scale User-Focused Dataset for…

200 papers

People increasingly use videos on the Web as a source for learning. To support this way of learning, researchers and developers are continuously developing tools, proposing guidelines, analyzing data, and conducting experiments. However, it…

Multimedia · Computer Science 2023-08-15 Evelyn Navarrete , Andreas Nehring , Sascha Schanze , Ralph Ewerth , Anett Hoppe

While generative models such as text-to-image, large language models and text-to-video have seen significant progress, the extension to text-to-virtual-reality remains largely unexplored, due to a deficit in training data and the complexity…

Computer Vision and Pattern Recognition · Computer Science 2025-03-12 Vriksha Srihari , R. Bhavya , Shruti Jayaraman , V. Mary Anita Rajam

The escalating quality of video generated by advanced video generation methods results in new security challenges, while there have been few relevant research efforts: 1) There is no open-source dataset for generated video detection, 2) No…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Long Ma , Zhiyuan Yan , Qinglang Guo , Yong Liao , Haiyang Yu , Pengyuan Zhou

Automatically generating scripts (i.e. sequences of key steps described in text) from video demonstrations and reasoning about the subsequent steps are crucial to the modern AI virtual assistants to guide humans to complete everyday tasks,…

Computation and Language · Computer Science 2024-01-22 Jingyuan Qi , Minqian Liu , Ying Shen , Zhiyang Xu , Lifu Huang

Among numerous videos shared on the web, well-edited ones always attract more attention. However, it is difficult for inexperienced users to make well-edited videos because it requires professional expertise and immense manual labor. To…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Yu Xiong , Fabian Caba Heilbron , Dahua Lin

While there exists a lot of work on explainable complaint mining, articulating user concerns through text or video remains a significant challenge, often leaving issues unresolved. Users frequently struggle to express their complaints…

Computer Vision and Pattern Recognition · Computer Science 2025-09-25 Sarmistha Das , R E Zera Marveen Lyngkhoi , Kirtan Jain , Vinayak Goyal , Sriparna Saha , Manish Gupta

In this paper, we initiate an attempt of developing an end-to-end chat-centric video understanding system, coined as VideoChat. It integrates video foundation models and large language models via a learnable neural interface, excelling in…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 KunChang Li , Yinan He , Yi Wang , Yizhuo Li , Wenhai Wang , Ping Luo , Yali Wang , Limin Wang , Yu Qiao

Spatio-temporal consistency is a critical research topic in video generation. A qualified generated video segment must ensure plot plausibility and coherence while maintaining visual consistency of objects and scenes across varying…

Computer Vision and Pattern Recognition · Computer Science 2025-03-11 Runze Zhang , Guoguang Du , Xiaochuan Li , Qi Jia , Liang Jin , Lu Liu , Jingjing Wang , Cong Xu , Zhenhua Guo , Yaqian Zhao , Xiaoli Gong , Rengang Li , Baoyu Fan

The goal of text-to-video retrieval is to search large databases for relevant videos based on text queries. Existing methods have progressed to handling explicit queries where the visual content of interest is described explicitly; however,…

Computer Vision and Pattern Recognition · Computer Science 2025-11-18 Yiqing Shen , Chenxiao Fan , Chenjia Li , Mathias Unberath

Multimodal summarization with multimodal output (MSMO) has emerged as a promising research direction. Nonetheless, numerous limitations exist within existing public MSMO datasets, including insufficient maintenance, data inaccessibility,…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Jielin Qiu , Jiacheng Zhu , William Han , Aditesh Kumar , Karthik Mittal , Claire Jin , Zhengyuan Yang , Linjie Li , Jianfeng Wang , Ding Zhao , Bo Li , Lijuan Wang

Generative AI (GenAI) tools enhance social media video creation by streamlining tasks such as scriptwriting, visual and audio generation, and editing. These tools enable the creation of new content, including text, images, audio, and video,…

Human-Computer Interaction · Computer Science 2025-03-06 Torin Anderson , Shuo Niu

Video-to-audio generation (V2A) is of increasing importance in domains such as film post-production, AR/VR, and sound design, particularly for the creation of Foley sound effects synchronized with on-screen actions. Foley requires…

Sound · Computer Science 2025-11-25 Satvik Dixit , Koichi Saito , Zhi Zhong , Yuki Mitsufuji , Chris Donahue

The generative AI revolution has recently expanded to videos. Nevertheless, current state-of-the-art video models are still lagging behind image models in terms of visual quality and user control over the generated content. In this work, we…

Computer Vision and Pattern Recognition · Computer Science 2023-11-21 Michal Geyer , Omer Bar-Tal , Shai Bagon , Tali Dekel

The training of controllable text-to-video (T2V) models relies heavily on the alignment between videos and captions, yet little existing research connects video caption evaluation with T2V generation assessment. This paper introduces…

Artificial Intelligence · Computer Science 2025-05-20 Xinlong Chen , Yuanxing Zhang , Chongling Rao , Yushuo Guan , Jiaheng Liu , Fuzheng Zhang , Chengru Song , Qiang Liu , Di Zhang , Tieniu Tan

We present a new large-scale multilingual video description dataset, VATEX, which contains over 41,250 videos and 825,000 captions in both English and Chinese. Among the captions, there are over 206,000 English-Chinese parallel translation…

Computer Vision and Pattern Recognition · Computer Science 2020-06-18 Xin Wang , Jiawei Wu , Junkun Chen , Lei Li , Yuan-Fang Wang , William Yang Wang

In this work, we present TalkCuts, a large-scale dataset designed to facilitate the study of multi-shot human speech video generation. Unlike existing datasets that focus on single-shot, static viewpoints, TalkCuts offers 164k clips…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Jiaben Chen , Zixin Wang , Ailing Zeng , Yang Fu , Xueyang Yu , Siyuan Cen , Julian Tanke , Yihang Chen , Koichi Saito , Yuki Mitsufuji , Chuang Gan

We present Emu Video, a text-to-video generation model that factorizes the generation into two steps: first generating an image conditioned on the text, and then generating a video conditioned on the text and the generated image. We…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Rohit Girdhar , Mannat Singh , Andrew Brown , Quentin Duval , Samaneh Azadi , Sai Saketh Rambhatla , Akbar Shah , Xi Yin , Devi Parikh , Ishan Misra

Video captioning automatically generates short descriptions of the video content, usually in form of a single sentence. Many methods have been proposed for solving this task. A large dataset called MSR Video to Text (MSR-VTT) is often used…

Computer Vision and Pattern Recognition · Computer Science 2024-02-27 Haoran Chen , Jianmin Li , Simone Frintrop , Xiaolin Hu

Automatic transcriptions of consumer-generated multi-media content such as "Youtube" videos still exhibit high word error rates. Such data typically occupies a very broad domain, has been recorded in challenging conditions, with cheap…

Computation and Language · Computer Science 2017-12-08 Abhinav Gupta , Yajie Miao , Leonardo Neves , Florian Metze

While current video generation focuses on text or image conditions, practical applications like video editing and vlogging often need to seamlessly connect separate clips. In our work, we introduce Video Connecting, an innovative task that…

Computer Vision and Pattern Recognition · Computer Science 2026-01-28 Zhiyu Yin , Zhipeng Liu , Kehai Chen , Lemao Liu , Jin Liu , Hong-Dong Li , Yang Xiang , Min Zhang
‹ Prev 1 4 5 6 7 8 10 Next ›