English
Related papers

Related papers: Generating Video Descriptions with Topic Guidance

200 papers

A big part of achieving Artificial General Intelligence(AGI) is to build a machine that can see and listen like humans. Much work has focused on designing models for image classification, video classification, object detection, pose…

Computer Vision and Pattern Recognition · Computer Science 2021-08-31 Ruotian Luo

Automatic video captioning is challenging due to the complex interactions in dynamic real scenes. A comprehensive system would ultimately localize and track the objects, actions and interactions present in a video and generate a description…

Computer Vision and Pattern Recognition · Computer Science 2016-10-19 Mihai Zanfir , Elisabeta Marinoiu , Cristian Sminchisescu

With the recent popularity of animated GIFs on social media, there is need for ways to index them with rich metadata. To advance research on animated GIF understanding, we collected a new dataset, Tumblr GIF (TGIF), with 100K animated GIFs…

Computer Vision and Pattern Recognition · Computer Science 2016-04-13 Yuncheng Li , Yale Song , Liangliang Cao , Joel Tetreault , Larry Goldberg , Alejandro Jaimes , Jiebo Luo

Generating descriptions for videos has many applications including assisting blind people and human-robot interaction. The recent advances in image captioning as well as the release of large-scale movie description datasets such as MPII…

Computer Vision and Pattern Recognition · Computer Science 2015-06-05 Anna Rohrbach , Marcus Rohrbach , Bernt Schiele

Video captioning which automatically translates video clips into natural language sentences is a very important task in computer vision. By virtue of recent deep learning technologies, e.g., convolutional neural networks (CNNs) and…

Computer Vision and Pattern Recognition · Computer Science 2016-11-18 Junbo Wang , Wei Wang , Yan Huang , Liang Wang , Tieniu Tan

The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant shortcomings, including low temporal consistency, poor-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Zhiyu Tan , Xiaomeng Yang , Luozheng Qin , Hao Li

With the rapid development of AI-generated content (AIGC), video generation has emerged as one of its most dynamic and impactful subfields. In particular, the advancement of video generation foundation models has led to growing demand for…

Automatic generation of video descriptions in natural language, also called video captioning, aims to understand the visual content of the video and produce a natural language sentence depicting the objects and actions in the scene. This…

Computer Vision and Pattern Recognition · Computer Science 2020-12-15 Begum Citamak , Ozan Caglayan , Menekse Kuyu , Erkut Erdem , Aykut Erdem , Pranava Madhyastha , Lucia Specia

Generative models have shown significant achievements in audio generation tasks. However, existing models struggle with complex and detailed prompts, leading to potential performance degradation. We hypothesize that this problem stems from…

The task of video captioning, that is, the automatic generation of sentences describing a sequence of actions in a video, has attracted an increasing attention recently. The complex and high-dimensional representation of video data makes it…

Computer Vision and Pattern Recognition · Computer Science 2020-10-13 Menatallh Hammad , May Hammad , Mohamed Elshenawy

In this work, we present Auto-captions on GIF, which is a new large-scale pre-training dataset for generic video understanding. All video-sentence pairs are created by automatically extracting and filtering video caption annotations from…

Computer Vision and Pattern Recognition · Computer Science 2020-07-07 Yingwei Pan , Yehao Li , Jianjie Luo , Jun Xu , Ting Yao , Tao Mei

Most existing multimodal machine translation (MMT) datasets are predominantly composed of static images or short video clips, lacking extensive video data across diverse domains and topics. As a result, they fail to meet the demands of…

Computation and Language · Computer Science 2025-05-12 Jinze Lv , Jian Chen , Zi Long , Xianghua Fu , Yin Chen

In light of recent advances in multimodal Large Language Models (LLMs), there is increasing attention to scaling them from image-text data to more informative real-world videos. Compared to static images, video poses unique challenges for…

Computer Vision and Pattern Recognition · Computer Science 2024-06-04 Yang Jin , Zhicheng Sun , Kun Xu , Kun Xu , Liwei Chen , Hao Jiang , Quzhe Huang , Chengru Song , Yuliang Liu , Di Zhang , Yang Song , Kun Gai , Yadong Mu

The automatic generation of representative natural language descriptions for observable patterns in time series data enhances interpretability, simplifies analysis and increases cross-domain utility of temporal data. While pre-trained…

Computation and Language · Computer Science 2025-01-06 Mohamed Trabelsi , Aidan Boyd , Jin Cao , Huseyin Uzunalioglu

Generative vision-language models can produce fluent medical image captions but remain prone to hallucination, over-specific diagnostic claims, and factual inconsistency-serious issues in pathology. We investigate retrieval-guided…

Computer Vision and Pattern Recognition · Computer Science 2026-05-05 Md. Enamul Hoq , Wataru Uegami , Saghir Alfasly , Ghazal Alabtah , Sahar Rahimi Malakshan , Armita Kazemi , Alex T. Schmitgen , Fred Prior , H. R. Tizhoosh

Motivated by the recent progress in generative models, we introduce a model that generates images from natural language descriptions. The proposed model iteratively draws patches on a canvas, while attending to the relevant words in the…

Machine Learning · Computer Science 2016-03-01 Elman Mansimov , Emilio Parisotto , Jimmy Lei Ba , Ruslan Salakhutdinov

We are creating multimedia contents everyday and everywhere. While automatic content generation has played a fundamental challenge to multimedia community for decades, recent advances of deep learning have made this problem feasible. For…

Computer Vision and Pattern Recognition · Computer Science 2018-04-24 Yingwei Pan , Zhaofan Qiu , Ting Yao , Houqiang Li , Tao Mei

Facilitated by deep neural networks, video recommendation systems have made significant advances. Existing video recommendation systems directly exploit features from different modalities (e.g., user personal data, user behavior data, video…

Information Retrieval · Computer Science 2020-10-27 Shi Pu , Yijiang He , Zheng Li , Mao Zheng

With the rapid advancement of text-conditioned Video Generation Models (VGMs), the quality of generated videos has significantly improved, bringing these models closer to functioning as ``*world simulators*'' and making real-world-level…

Artificial Intelligence · Computer Science 2025-04-22 Haotong Yang , Qingyuan Zheng , Yunjian Gao , Yongkun Yang , Yangbo He , Zhouchen Lin , Muhan Zhang

Generating a video given the first several static frames is challenging as it anticipates reasonable future frames with temporal coherence. Besides video prediction, the ability to rewind from the last frame or infilling between the head…

Computer Vision and Pattern Recognition · Computer Science 2023-03-23 Tsu-Jui Fu , Licheng Yu , Ning Zhang , Cheng-Yang Fu , Jong-Chyi Su , William Yang Wang , Sean Bell
‹ Prev 1 4 5 6 7 8 10 Next ›