English
Related papers

Related papers: A multimodal approach for multi-label movie genre …

200 papers

Inspired by the remarkable advances in video analytics, research teams are stepping towards a greater ambition -- movie understanding. However, compared to those activity videos in conventional datasets, movies are significantly different.…

Computer Vision and Pattern Recognition · Computer Science 2019-10-25 Yu Xiong , Qingqiu Huang , Lingfeng Guo , Hang Zhou , Bolei Zhou , Dahua Lin

Video classification and analysis is always a popular and challenging field in computer vision. It is more than just simple image classification due to the correlation with respect to the semantic contents of subsequent frames brings…

Computer Vision and Pattern Recognition · Computer Science 2021-03-30 Yilin Wang , Jiayi Ye

We introduce TV show Retrieval (TVR), a new multimodal retrieval dataset. TVR requires systems to understand both videos and their associated subtitle (dialogue) texts, making it more realistic. The dataset contains 109K queries collected…

Computer Vision and Pattern Recognition · Computer Science 2020-08-19 Jie Lei , Licheng Yu , Tamara L. Berg , Mohit Bansal

In this work, we explore different approaches to combine modalities for the problem of automated age-suitability rating of movie trailers. First, we introduce a new dataset containing videos of movie trailers in English downloaded from IMDB…

Machine Learning · Computer Science 2021-01-29 Mahsa Shafaei , Christos Smailis , Ioannis A. Kakadiaris , Thamar Solorio

Characterizing and quantifying gender representation disparities in audiovisual storytelling contents is necessary to grasp how stereotypes may perpetuate on screen. In this article, we consider the high-level construct of objectification…

Social tagging of movies reveals a wide range of heterogeneous information about movies, like the genre, plot structure, soundtracks, metadata, visual and emotional experiences. Such information can be valuable in building automatic systems…

Computation and Language · Computer Science 2018-02-26 Sudipta Kar , Suraj Maharjan , A. Pastor López-Monroy , Thamar Solorio

Story video-text alignment, a core task in computational story understanding, aims to align video clips with corresponding sentences in their descriptions. However, progress on the task has been held back by the scarcity of manually…

Computation and Language · Computer Science 2024-10-04 Yidan Sun , Jianfei Yu , Boyang Li

In this study, we investigated multi-modal approaches using images, descriptions, and titles to categorize e-commerce products on Amazon. Specifically, we examined late fusion models, where the modalities are fused at the decision level.…

Machine Learning · Computer Science 2019-09-18 Pasawee Wirojwatanakul , Artit Wangperawong

Robust face clustering is a vital step in enabling computational understanding of visual character portrayal in media. Face clustering for long-form content is challenging because of variations in appearance and lack of supporting…

Computer Vision and Pattern Recognition · Computer Science 2022-03-01 Krishna Somandepalli , Rajat Hebbar , Shrikanth Narayanan

Our objective in this work is long range understanding of the narrative structure of movies. Instead of considering the entire movie, we propose to learn from the `key scenes' of the movie, providing a condensed look at the full storyline.…

Computer Vision and Pattern Recognition · Computer Science 2020-10-26 Max Bain , Arsha Nagrani , Andrew Brown , Andrew Zisserman

Music genre classification has been widely studied in past few years for its various applications in music information retrieval. Previous works tend to perform unsatisfactorily, since those methods only use audio content or jointly use…

Sound · Computer Science 2023-06-13 Ganghui Ru , Xulong Zhang , Jianzong Wang , Ning Cheng , Jing Xiao

Massive multi-modality datasets play a significant role in facilitating the success of large video-language models. However, current video-language datasets primarily provide text descriptions for visual frames, considering audio to be…

The goal of this study is to develop and analyze multimodal models for predicting experienced affective responses of viewers watching movie clips. We develop hybrid multimodal prediction models based on both the video and audio of the…

Computer Vision and Pattern Recognition · Computer Science 2019-09-18 Ha Thi Phuong Thao , Dorien Herremans , Gemma Roig

Multi-label image classification is a foundational topic in various domains. Multimodal learning approaches have recently achieved outstanding results in image representation and single-label image classification. For instance, Contrastive…

Computer Vision and Pattern Recognition · Computer Science 2022-11-11 Fengjun Wang , Sarai Mizrachi , Moran Beladev , Guy Nadav , Gil Amsalem , Karen Lastmann Assaraf , Hadas Harush Boker

We propose a new model for speaker naming in movies that leverages visual, textual, and acoustic modalities in an unified optimization framework. To evaluate the performance of our model, we introduce a new dataset consisting of six…

Computation and Language · Computer Science 2018-09-25 Mahmoud Azab , Mingzhe Wang , Max Smith , Noriyuki Kojima , Jia Deng , Rada Mihalcea

This paper considers the problem of characterizing stories by inferring properties such as theme and style using written synopses and reviews of movies. We experiment with a multi-label dataset of movie synopses and a tagset representing…

Computation and Language · Computer Science 2020-10-12 Sudipta Kar , Gustavo Aguilar , Mirella Lapata , Thamar Solorio

Previous research has demonstrated the advantages of integrating data from multiple sources over traditional unimodal data, leading to the emergence of numerous novel multimodal applications. We propose a multimodal classification benchmark…

Machine Learning · Computer Science 2023-12-20 Jiaying Lu , Yongchen Qian , Shifan Zhao , Yuanzhe Xi , Carl Yang

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of…

Computer Vision and Pattern Recognition · Computer Science 2020-04-29 Anyi Rao , Linning Xu , Yu Xiong , Guodong Xu , Qingqiu Huang , Bolei Zhou , Dahua Lin

As the volume of digital image data increases, the effectiveness of image classification intensifies. This study introduces a robust multi-label classification system designed to assign multiple labels to a single image, addressing the…

Computer Vision and Pattern Recognition · Computer Science 2025-01-06 Haixu Liu , Penghao Jiang , Zerui Tao

Oscar nominations are an important factor in the movie industry because they can boost both the visibility and the commercial success. This work explores whether it is possible to predict Oscar nominations for screenplays using modern…

Information Retrieval · Computer Science 2025-11-11 Francis Gross