English
Related papers

Related papers: The MeLa BitChute Dataset

200 papers

Much of the delivery of University education is now by synchronous or asynchronous video. For students, one of the challenges is managing the sheer volume of such video material as video presentations of taught material are difficult to…

Multimedia · Computer Science 2021-06-28 Hyowon Lee , Mingming Liu , Michael Scriney , Alan F. Smeaton

EventNet is a large-scale video corpus and event ontology consisting of 500 events associated with event-specific concepts. In order to improve the quality of the current EventNet, we conduct the following steps and introduce EventNet…

Computer Vision and Pattern Recognition · Computer Science 2016-09-13 Dongang Wang , Zheng Shou , Hongyi Liu , Shih-Fu Chang

As short videos have become the primary form of content consumption across various industries, accurately predicting their popularity has become key to enhancing user engagement and optimizing business strategies. This report presents a…

Multimedia · Computer Science 2025-02-25 Jiacheng Lu , Mingyuan Xiao , Weijian Wang , Yuxin Du , Zhengze Wu , Cheng Hua

The quality of the data and annotation upper-bounds the quality of a downstream model. While there exist large text corpora and image-text pairs, high-quality video-text data is much harder to collect. First of all, manual labeling is more…

Telegram has emerged as a major platform for large-scale video piracy, where copyrighted content is rapidly distributed among users. Despite its prominence, the structural and operational dynamics of this ecosystem remain insufficiently…

Cryptography and Security · Computer Science 2026-05-21 Sadikshya Gyawali , Jaishnoor Kaur , Taylor Graham , Josef Horacek , Nowshin Tabassum , Shirin Nilizadeh , Sayak Saha Roy

The profusion of user generated content caused by the rise of social media platforms has enabled a surge in research relating to fields such as information retrieval, recommender systems, data mining and machine learning. However, the lack…

Social and Information Networks · Computer Science 2018-01-23 Nuno Moniz , Luís Torgo

We present a new dataset with annotated eye movements. The dataset consists of over 800,000 gaze points recorded during a car ride in the real world and in the simulator. In total, the eye movements of 19 subjects were annotated. In this…

Computer Vision and Pattern Recognition · Computer Science 2021-01-13 Wolfgang Fuhl , Enkelejda Kasneci

Short video platforms like TikTok, Instagram Reels, and YouTube Shorts have gained immense popularity in the last few years and are responsible for a large and growing fraction of Internet traffic. We identify two unique opportunities for…

Networking and Internet Architecture · Computer Science 2026-05-07 Maleeha Masood , Shreya Kannan , Om Chabra , Deepak Vasisht , Indranil Gupta

Egocentric video has seen increased interest in recent years, as it is used in a range of areas. However, most existing datasets are limited to a single perspective. In this paper, we present the CASTLE 2024 dataset, a multimodal collection…

The Multi-language Video Subtitle Dataset is a comprehensive collection designed to support research in text recognition across multiple languages. This dataset includes 4,224 subtitle images extracted from 24 videos sourced from online…

Computer Vision and Pattern Recognition · Computer Science 2024-11-11 Thanadol Singkhornart , Olarik Surinta

Recent developments in vision-language models have significantly advanced video understanding. Existing datasets and tasks, however, have notable limitations. Most datasets are confined to short videos with limited events and narrow…

Computer Vision and Pattern Recognition · Computer Science 2025-01-13 Ridouane Ghermi , Xi Wang , Vicky Kalogeiton , Ivan Laptev

Text-to-video generative models convert textual prompts into dynamic visual content, offering wide-ranging applications in film production, gaming, and education. However, their real-world performance often falls short of user expectations.…

Computer Vision and Pattern Recognition · Computer Science 2025-05-14 Wenhao Wang , Yi Yang

This paper introduces ChinaOpen, a dataset sourced from Bilibili, a popular Chinese video-sharing website, for open-world multimodal learning. While the state-of-the-art multimodal learning networks have shown impressive performance in…

Multimedia · Computer Science 2023-08-08 Aozhu Chen , Ziyuan Wang , Chengbo Dong , Kaibin Tian , Ruixiang Zhao , Xun Liang , Zhanhui Kang , Xirong Li

Telegram is one of the most popular instant messaging apps in today's digital age. In addition to providing a private messaging service, Telegram, with its channels, represents a valid medium for rapidly broadcasting content to a large…

Computers and Society · Computer Science 2025-03-04 Massimo La Morgia , Alessandro Mei , Alberto Maria Mongardini

The visual world around us constantly evolves, from real-time news and social media trends to global infrastructure changes visible through satellite imagery and augmented reality enhancements. However, Multimodal Large Language Models…

Computer Vision and Pattern Recognition · Computer Science 2025-07-02 Mingyang Fu , Yuyang Peng , Dongping Chen , Zetong Zhou , Benlin Liu , Yao Wan , Zhou Zhao , Philip S. Yu , Ranjay Krishna

Food image segmentation is a critical task for dietary analysis, enabling accurate estimation of food volume and nutrients. However, current methods suffer from limited multi-view data and poor generalization to new viewpoints. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2026-01-21 Ahmad AlMughrabi , Guillermo Rivo , Carlos Jiménez-Farfán , Umair Haroon , Farid Al-Areqi , Hyunjun Jung , Benjamin Busam , Ricardo Marques , Petia Radeva

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

In this work, we introduce a new problem, named as {\em story-preserving long video truncation}, that requires an algorithm to automatically truncate a long-duration video into multiple short and attractive sub-videos with each one…

Computer Vision and Pattern Recognition · Computer Science 2019-10-15 Fan Yang , Xiao Liu , Dongliang He , Chuang Gan , Jian Wang , Chao Li , Fu Li , Shilei Wen

We publicly release a new large-scale dataset, called SearchQA, for machine comprehension, or question-answering. Unlike recently released datasets, such as DeepMind CNN/DailyMail and SQuAD, the proposed SearchQA was constructed to reflect…

Computation and Language · Computer Science 2017-06-13 Matthew Dunn , Levent Sagun , Mike Higgins , V. Ugur Guney , Volkan Cirik , Kyunghyun Cho

This paper investigates the use of automatically collected web audio data for the task of spoken language recognition. We generate semi-random search phrases from language-specific Wikipedia data that are then used to retrieve videos from…

Audio and Speech Processing · Electrical Eng. & Systems 2020-11-30 Jörgen Valk , Tanel Alumäe
‹ Prev 1 8 9 10 Next ›