English
Related papers

Related papers: Multimodal Feature Fusion for Video Advertisements…

200 papers

This paper addresses the question of emotion classification. The task consists in predicting emotion labels (taken among a set of possible labels) best describing the emotions contained in short video clips. Building on a standard framework…

Computer Vision and Pattern Recognition · Computer Science 2017-09-22 Valentin Vielzeuf , Stéphane Pateux , Frédéric Jurie

Video summarization, by selecting the most informative and/or user-relevant parts of original videos to create concise summary videos, has high research value and consumer demand in today's video proliferation era. Multi-modal video…

Computer Vision and Pattern Recognition · Computer Science 2025-03-04 Yaowei Guo , Jiazheng Xing , Xiaojun Hou , Shuo Xin , Juntao Jiang , Demetri Terzopoulos , Chenfanfu Jiang , Yong Liu

The application of video captioning models aims at translating the content of videos by using accurate natural language. Due to the complex nature inbetween object interaction in the video, the comprehensive understanding of spatio-temporal…

Computer Vision and Pattern Recognition · Computer Science 2023-08-15 Yutao Jin , Bin Liu , Jing Wang

Compared to images, videos better reflect real-world acquisition and possess valuable temporal cues. However, existing multi-sensor fusion research predominantly integrates complementary context from multiple images rather than videos due…

Computer Vision and Pattern Recognition · Computer Science 2026-03-03 Linfeng Tang , Yeda Wang , Meiqi Gong , Zizhuo Li , Yuxin Deng , Xunpeng Yi , Chunyu Li , Han Xu , Hao Zhang , Jiayi Ma

Feature alignment serves as the primary mechanism for fusing multimodal data. We put forth a feature alignment approach that achieves full integration of multimodal information. This is accomplished via an alternating process of shifting…

Computer Vision and Pattern Recognition · Computer Science 2024-06-14 Jiahao Qin

This thesis presents an innovative approach to automate video thumbnail selection for traditional broadcast content. Our methodology establishes stringent criteria for diverse, representative, and aesthetically pleasing thumbnails,…

Computer Vision and Pattern Recognition · Computer Science 2024-10-29 Elia Fantini

Multimodal learning integrates data from diverse sensors to effectively harness information from different modalities. However, recent studies reveal that joint learning often overfits certain modalities while neglecting others, leading to…

Computer Vision and Pattern Recognition · Computer Science 2026-03-11 Feng Yu , Xiangyu Wu , Yang Yang , Jianfeng Lu

Gesture recognition is a much studied research area which has myriad real-world applications including robotics and human-machine interaction. Current gesture recognition methods have focused on recognising isolated gestures, and existing…

Computer Vision and Pattern Recognition · Computer Science 2021-09-22 Harshala Gammulle , Simon Denman , Sridha Sridharan , Clinton Fookes

Exploiting both audio and visual modalities for video classification is a challenging task, as the existing methods require large model architectures, leading to high computational complexity and resource requirements. Smaller…

Computer Vision and Pattern Recognition · Computer Science 2024-08-27 Mahrukh Awan , Asmar Nadeem , Muhammad Junaid Awan , Armin Mustafa , Syed Sameed Husain

Multimodal image fusion and semantic segmentation are critical for autonomous driving. Despite advancements, current models often struggle with segmenting densely packed elements due to a lack of comprehensive fusion features for guidance…

Computer Vision and Pattern Recognition · Computer Science 2025-06-25 Daixun Li , Weiying Xie , Mingxiang Cao , Yunke Wang , Yusi Zhang , Leyuan Fang , Yunsong Li , Chang Xu

Multimodal deep learning systems which employ multiple modalities like text, image, audio, video, etc., are showing better performance in comparison with individual modalities (i.e., unimodal) systems. Multimodal machine learning involves…

Machine Learning · Computer Science 2022-01-19 Anil Rahate , Rahee Walambe , Sheela Ramanna , Ketan Kotecha

Modern video summarization methods are based on deep neural networks that require a large amount of annotated data for training. However, existing datasets for video summarization are small-scale, easily leading to over-fitting of the deep…

Computer Vision and Pattern Recognition · Computer Science 2022-10-20 Li Haopeng , Ke Qiuhong , Gong Mingming , Tom Drummond

Effective fusion of data from multiple modalities, such as video, speech, and text, is challenging due to the heterogeneous nature of multimodal data. In this paper, we propose adaptive fusion techniques that aim to model context from…

Computation and Language · Computer Science 2021-01-27 Gaurav Sahu , Olga Vechtomova

We propose a tri-modal architecture to predict Big Five personality trait scores from video clips with different channels for audio, text, and video data. For each channel, stacked Convolutional Neural Networks are employed. The channels…

Artificial Intelligence · Computer Science 2018-05-17 Onno Kampman , Elham J. Barezi , Dario Bertero , Pascale Fung

The rapid rise of video content on platforms such as TikTok and YouTube has transformed information dissemination, but it has also facilitated the spread of harmful content, particularly hate videos. Despite significant efforts to combat…

Multimedia · Computer Science 2025-05-20 Yinghui Zhang , Tailin Chen , Yuchen Zhang , Zeyu Fu

Video summarization methods are usually classified into shot-level or frame-level methods, which are individually used in a general way. This paper investigates the underlying complementarity between the frame-level and shot-level methods,…

Computer Vision and Pattern Recognition · Computer Science 2022-07-05 Yubo An , Shenghui Zhao , Guoqiang Zhang

Multimodal video understanding plays a crucial role in tasks such as action recognition and emotion classification by combining information from different modalities. However, multimodal models are prone to overfitting strong modalities,…

Computer Vision and Pattern Recognition · Computer Science 2025-10-14 Xiaoyu Ma , Ding Ding , Hao Chen

This research introduces a multimodal system designed to detect fraud and fare evasion in public transportation by analyzing closed circuit television (CCTV) and audio data. The proposed solution uses the Vision Transformer for Video…

Software Engineering · Computer Science 2025-10-03 Peter Wauyo , Dalia Bwiza , Alain Murara , Edwin Mugume , Eric Umuhoza

Finding relevant moments and highlights in videos according to natural language queries is a natural and highly valuable common need in the current video content explosion era. Nevertheless, jointly conducting moment retrieval and highlight…

Computer Vision and Pattern Recognition · Computer Science 2022-03-29 Ye Liu , Siyuan Li , Yang Wu , Chang Wen Chen , Ying Shan , Xiaohu Qie

Over the past decade, the evolution of video-sharing platforms has attracted a significant amount of investments on contextual advertising. The common contextual advertising platforms utilize the information provided by users to integrate…

Computer Vision and Pattern Recognition · Computer Science 2020-06-29 Ivan Bacher , Hossein Javidnia , Soumyabrata Dev , Rahul Agrahari , Murhaf Hossari , Matthew Nicholson , Clare Conran , Jian Tang , Peng Song , David Corrigan , François Pitié