中文
相关论文

相关论文: MovieNet: A Holistic Dataset for Movie Understandi…

200 篇论文

Movie screenplay summarization is challenging, as it requires an understanding of long input contexts and various elements unique to movies. Large language models have shown significant advancements in document summarization, but they often…

计算与语言 · 计算机科学 2024-08-13 Rohit Saxena , Frank Keller

Despite recent advances of AI, story understanding remains an open and under-investigated problem. We collect, preprocess, and publicly release a video-language story dataset, Synopses of Movie Narratives (SyMoN), containing 5,193 video…

计算机视觉与模式识别 · 计算机科学 2023-04-06 Yidan Sun , Qin Chao , Yangfeng Ji , Boyang Li

Longform media such as movies have complex narrative structures, with events spanning a rich variety of ambient visual scenes. Domain specific challenges associated with visual scenes in movies include transitions, person coverage, and a…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Digbalay Bose , Rajat Hebbar , Krishna Somandepalli , Haoyang Zhang , Yin Cui , Kree Cole-McLaughlin , Huisheng Wang , Shrikanth Narayanan

Film media is a rich form of artistic expression. Unlike photography, and short videos, movies contain a storyline that is deliberately complex and intricate in order to engage its audience. In this paper we present a large scale study…

计算机视觉与模式识别 · 计算机科学 2019-08-09 Paola Cascante-Bonilla , Kalpathy Sitaraman , Mengjia Luo , Vicente Ordonez

We introduce the MovieQA dataset which aims to evaluate automatic story comprehension from both video and text. The dataset consists of 14,944 questions about 408 movies with high semantic diversity. The questions range from simpler "Who"…

计算机视觉与模式识别 · 计算机科学 2016-09-22 Makarand Tapaswi , Yukun Zhu , Rainer Stiefelhagen , Antonio Torralba , Raquel Urtasun , Sanja Fidler

Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames…

计算机视觉与模式识别 · 计算机科学 2024-10-22 Ruchit Rawal , Khalid Saifullah , Miquel Farré , Ronen Basri , David Jacobs , Gowthami Somepalli , Tom Goldstein

This paper introduces HarmonySet, a comprehensive dataset designed to advance video-music understanding. HarmonySet consists of 48,328 diverse video-music pairs, annotated with detailed information on rhythmic synchronization, emotional…

计算机视觉与模式识别 · 计算机科学 2025-03-05 Zitang Zhou , Ke Mei , Yu Lu , Tianyi Wang , Fengyun Rao

Understanding movies and their structural patterns is a crucial task in decoding the craft of video editing. While previous works have developed tools for general analysis, such as detecting characters or recognizing cinematography…

计算机视觉与模式识别 · 计算机科学 2022-10-25 Alejandro Pardo , Fabian Caba Heilbron , Juan León Alcázar , Ali Thabet , Bernard Ghanem

Our objective in this work is long range understanding of the narrative structure of movies. Instead of considering the entire movie, we propose to learn from the `key scenes' of the movie, providing a condensed look at the full storyline.…

计算机视觉与模式识别 · 计算机科学 2020-10-26 Max Bain , Arsha Nagrani , Andrew Brown , Andrew Zisserman

This paper introduces VideoMind, a video-centric omni-modal dataset designed for deep video content cognition and enhanced multi-modal feature representation. The dataset comprises 103K video samples (3K reserved for testing), each paired…

计算机视觉与模式识别 · 计算机科学 2025-07-25 Baoyao Yang , Wanyun Li , Dixin Chen , Junxiang Chen , Wenbin Yao , Haifeng Lin

Monitoring animal behavior can facilitate conservation efforts by providing key insights into wildlife health, population status, and ecosystem function. Automatic recognition of animals and their behaviors is critical for capitalizing on…

计算机视觉与模式识别 · 计算机科学 2023-06-02 Jun Chen , Ming Hu , Darren J. Coker , Michael L. Berumen , Blair Costelloe , Sara Beery , Anna Rohrbach , Mohamed Elhoseiny

Discovering content and stories in movies is one of the most important concepts in multimedia content research studies. Network models have proven to be an efficient choice for this purpose. When an audience watches a movie, they usually…

机器学习 · 计算机科学 2019-10-22 Youssef Mourchid , Benjamin Renoust , Olivier Roupin , Le Van , Hocine Cherifi , Mohammed El Hassouni

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online…

计算机视觉与模式识别 · 计算机科学 2024-11-11 Thanadol Singkhornart , Olarik Surinta

Recent developments in vision-language models have significantly advanced video understanding. Existing datasets and tasks, however, have notable limitations. Most datasets are confined to short videos with limited events and narrow…

计算机视觉与模式识别 · 计算机科学 2025-01-13 Ridouane Ghermi , Xi Wang , Vicky Kalogeiton , Ivan Laptev

Descriptive video service (DVS) provides linguistic descriptions of movies and allows visually impaired people to follow a movie along with their peers. Such descriptions are by design mainly visual and thus naturally form an interesting…

计算机视觉与模式识别 · 计算机科学 2015-01-13 Anna Rohrbach , Marcus Rohrbach , Niket Tandon , Bernt Schiele

Film, a classic image style, is culturally significant to the whole photographic industry since it marks the birth of photography. However, film photography is time-consuming and expensive, necessitating a more efficient method for…

计算机视觉与模式识别 · 计算机科学 2023-11-06 Zinuo Li , Xuhang Chen , Shuqiang Wang , Chi-Man Pun

Inspired by the remarkable advances in video analytics, research teams are stepping towards a greater ambition -- movie understanding. However, compared to those activity videos in conventional datasets, movies are significantly different.…

计算机视觉与模式识别 · 计算机科学 2019-10-25 Yu Xiong , Qingqiu Huang , Lingfeng Guo , Hang Zhou , Bolei Zhou , Dahua Lin

Scene, as the crucial unit of storytelling in movies, contains complex activities of actors and their interactions in a physical environment. Identifying the composition of scenes serves as a critical step towards semantic understanding of…

计算机视觉与模式识别 · 计算机科学 2020-04-29 Anyi Rao , Linning Xu , Yu Xiong , Guodong Xu , Qingqiu Huang , Bolei Zhou , Dahua Lin

Despite the significant impact of visual events on human cognition, understanding events in videos remains a challenging task for AI due to their complex structures, semantic hierarchies, and dynamic evolution. To address this, we propose…

计算机视觉与模式识别 · 计算机科学 2025-06-04 Baoyu Liang , Qile Su , Shoutai Zhu , Yuchen Liang , Chao Tong

Video recognition has been advanced in recent years by benchmarks with rich annotations. However, research is still mainly limited to human action or sports recognition - focusing on a highly specific video understanding task and thus…

计算机视觉与模式识别 · 计算机科学 2020-12-16 Ali Diba , Mohsen Fayyaz , Vivek Sharma , Manohar Paluri , Jurgen Gall , Rainer Stiefelhagen , Luc Van Gool
‹ 上一页 1 2 3 10 下一页 ›