English
Related papers

Related papers: Multimodal Feature Fusion for Video Advertisements…

200 papers

Multimodal headline utilizes both video frames and transcripts to generate the natural language title of the videos. Due to a lack of large-scale, manually annotated data, the task of annotating grounded headlines for video is labor…

Computer Vision and Pattern Recognition · Computer Science 2022-11-15 Lingfeng Qiao , Chen Wu , Ye Liu , Haoyuan Peng , Di Yin , Bo Ren

We address the problem of text-guided video temporal grounding, which aims to identify the time interval of a certain event based on a natural language description. Different from most existing methods that only consider RGB images as…

Computer Vision and Pattern Recognition · Computer Science 2021-11-01 Yi-Wen Chen , Yi-Hsuan Tsai , Ming-Hsuan Yang

This technical report presents the 3rd winning solution for MTVG, a new task introduced in the 4-th Person in Context (PIC) Challenge at ACM MM 2022. MTVG aims at localizing the temporal boundary of the step in an untrimmed video based on a…

Computer Vision and Pattern Recognition · Computer Science 2022-08-15 Xiujun Shu , Wei Wen , Taian Guo , Sunan He , Chen Wu , Ruizhi Qiao

Deep learning-based methods have achieved promising results on surgical instrument segmentation. However, the high computation cost may limit the application of deep models to time-sensitive tasks such as online surgical video analysis for…

Computer Vision and Pattern Recognition · Computer Science 2021-07-27 Shan Lin , Fangbo Qin , Haonan Peng , Randall A. Bly , Kris S. Moe , Blake Hannaford

The topic diversity of open-domain videos leads to various vocabularies and linguistic expressions in describing video contents, and therefore, makes the video captioning task even more challenging. In this paper, we propose an unified…

Computer Vision and Pattern Recognition · Computer Science 2023-02-15 Shizhe Chen , Jia Chen , Qin Jin , Alexander Hauptmann

Assessing aesthetic preference is a fundamental task related to human cognition. It can also contribute to various practical applications such as image creation for online advertisements. Despite crucial influences of image quality,…

Machine Learning · Computer Science 2019-10-08 Kyung-Wha Park , JungHoon Lee , Sunyoung Kwon , Jung-Woo Ha , Kyung-Min Kim , Byoung-Tak Zhang

In this paper, the main task we aim to tackle is the multi-instance semi-supervised video object segmentation across a sequence of frames where only the first-frame box-level ground-truth is provided. Detection-based algorithms are widely…

Computer Vision and Pattern Recognition · Computer Science 2020-04-17 Mingjie Sun , Jimin Xiao , Eng Gee Lim , Bingfeng Zhang , Yao Zhao

Event classification is inherently sequential and multimodal. Therefore, deep neural models need to dynamically focus on the most relevant time window and/or modality of a video. In this study, we propose the Multi-level Attention Fusion…

Computer Vision and Pattern Recognition · Computer Science 2021-06-15 Mathilde Brousmiche , Jean Rouat , Stéphane Dupont

Video captioning is a popular task that challenges models to describe events in videos using natural language. In this work, we investigate the ability of various visual feature representations derived from state-of-the-art convolutional…

Computer Vision and Pattern Recognition · Computer Science 2021-01-18 Praveen S , Akhilesh Bharadwaj , Harsh Raj , Janhavi Dadhania , Ganesh Samarth C. A , Nikhil Pareek , S R M Prasanna

Traditional recommender systems heavily rely on ID features, which often encounter challenges related to cold-start and generalization. Modeling pre-extracted content features can mitigate these issues, but is still a suboptimal solution…

Information Retrieval · Computer Science 2024-04-10 Xiuqi Deng , Lu Xu , Xiyao Li , Jinkai Yu , Erpeng Xue , Zhongyuan Wang , Di Zhang , Zhaojie Liu , Guorui Zhou , Yang Song , Na Mou , Shen Jiang , Han Li

This study tackles the challenge of image matching in difficult scenarios, such as scenes with significant variations or limited texture, with a strong emphasis on computational efficiency. Previous studies have attempted to address this…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Khang Truong Giang , Soohwan Song , Sungho Jo

There is a perennial need in the online advertising industry to refresh ad creatives, i.e., images and text used for enticing online users towards a brand. Such refreshes are required to reduce the likelihood of ad fatigue among online…

Computation and Language · Computer Science 2020-08-05 Yichao Zhou , Shaunak Mishra , Manisha Verma , Narayan Bhamidipati , Wei Wang

This paper presents a novel deep neural network (DNN) for multimodal fusion of audio, video and text modalities for emotion recognition. The proposed DNN architecture has independent and shared layers which aim to learn the representation…

Computer Vision and Pattern Recognition · Computer Science 2019-07-09 Juan D. S. Ortega , Mohammed Senoussaoui , Eric Granger , Marco Pedersoli , Patrick Cardinal , Alessandro L. Koerich

Multi-modal learning has emerged as a crucial research direction, as integrating textual and visual information can substantially enhance performance in tasks such as classification, retrieval, and scene understanding. Despite advances with…

Computer Vision and Pattern Recognition · Computer Science 2025-07-31 Md. Mithun Hossain , Md. Shakil Hossain , Sudipto Chaki , M. F. Mridha

Video-based ads are a vital medium for brands to engage consumers, with social media platforms leveraging user data to optimize ad delivery and boost engagement. A crucial but under-explored aspect is the 'hooking period', the first three…

Multimedia · Computer Science 2026-02-27 Kunpeng Zhang , Poppy Zhang , Shawndra Hill , Amel Awadelkarim

In the last decade, video blogs (vlogs) have become an extremely popular method through which people express sentiment. The ubiquitousness of these videos has increased the importance of multimodal fusion models, which incorporate video and…

Computer Vision and Pattern Recognition · Computer Science 2018-07-04 Nathaniel Blanchard , Daniel Moreira , Aparna Bharati , Walter J. Scheirer

Current video retrieval systems, especially those used in competitions, primarily focus on querying individual keyframes or images rather than encoding an entire clip or video segment. However, queries often describe an action or event over…

Computer Vision and Pattern Recognition · Computer Science 2026-04-29 Quoc-Bao Nguyen-Le , Thanh-Huy Le-Nguyen

Dividing ads ranking system into retrieval, early, and final stages is a common practice in large scale ads recommendation to balance the efficiency and accuracy. The early stage ranking often uses efficient models to generate candidates…

Information Retrieval · Computer Science 2023-07-24 Xuewei Wang , Qiang Jin , Shengyu Huang , Min Zhang , Xi Liu , Zhengli Zhao , Yukun Chen , Zhengyu Zhang , Jiyan Yang , Ellie Wen , Sagar Chordia , Wenlin Chen , Qin Huang

Multimodal models have achieved remarkable success in natural image segmentation, yet they often underperform when applied to the medical domain. Through extensive study, we attribute this performance gap to the challenges of multimodal…

Computer Vision and Pattern Recognition · Computer Science 2025-09-11 Wenjun Yu , Yinchen Zhou , Jia-Xuan Jiang , Shubin Zeng , Yuee Li , Zhong Wang

This paper proposes a novel multimodal fusion approach, aiming to produce best possible decisions by integrating information coming from multiple media. While most of the past multimodal approaches either work by projecting the features of…

Artificial Intelligence · Computer Science 2018-08-23 Valentin Vielzeuf , Alexis Lechervy , Stéphane Pateux , Frédéric Jurie