English
Related papers

Related papers: The MeLa BitChute Dataset

200 papers

Current datasets for long-form video understanding often fall short of providing genuine long-form comprehension challenges, as many tasks derived from these datasets can be successfully tackled by analyzing just one or a few random frames…

Computer Vision and Pattern Recognition · Computer Science 2024-10-22 Ruchit Rawal , Khalid Saifullah , Miquel Farré , Ronen Basri , David Jacobs , Gowthami Somepalli , Tom Goldstein

We propose a novel framework for predicting the factuality of reporting of news media outlets by studying the user attention cycles in their YouTube channels. In particular, we design a rich set of features derived from the temporal…

Computation and Language · Computer Science 2021-08-31 Krasimira Bozhanova , Yoan Dinkov , Ivan Koychev , Maria Castaldo , Tommaso Venturini , Preslav Nakov

For research results to be comparable, it is important to have common datasets for experimentation and evaluation. The size of such datasets, however, can be an obstacle to their use. The Vimeo Creative Commons Collection (V3C) is a video…

Multimedia · Computer Science 2021-05-05 Luca Rossetto , Klaus Schoeffmann , Abraham Bernstein

Large-scale datasets have played a significant role in progress of neural network and deep learning areas. YouTube-8M is such a benchmark dataset for general multi-label video classification. It was created from over 7 million YouTube…

Machine Learning · Statistics 2017-06-27 Zhenzhen Zhong , Shujiao Huang , Cheng Zhan , Licheng Zhang , Zhiwei Xiao , Chang-Chun Wang , Pei Yang

In this paper, we introduce YTCommentVerse, a large-scale multilingual and multi-category dataset of YouTube comments. It contains over 32 million comments from 178,000 videos contributed by more than 20 million unique users spanning 15…

Social and Information Networks · Computer Science 2025-09-16 Hridoy Sankar Dutta , Biswadeep Khan

The share of videos in the internet traffic has been growing, therefore understanding how videos capture attention on a global scale is also of growing importance. Most current research focus on modeling the number of views, but we argue…

Social and Information Networks · Computer Science 2018-04-12 Siqi Wu , Marian-Andrei Rizoiu , Lexing Xie

In this paper, we present an updated version of the NELA-GT-2018 dataset (N{\o}rregaard, Horne, and Adal{\i} 2019), entitled NELA-GT-2019. NELA-GT-2019 contains 1.12M news articles from 260 sources collected between January 1st 2019 and…

Computers and Society · Computer Science 2020-03-30 Maurício Gruppi , Benjamin D. Horne , Sibel Adalı

Messaging platforms, especially those with a mobile focus, have become increasingly ubiquitous in society. These mobile messaging platforms can have deceivingly large user bases, and in addition to being a way for people to stay in touch,…

Social and Information Networks · Computer Science 2020-01-24 Jason Baumgartner , Savvas Zannettou , Megan Squire , Jeremy Blackburn

Established in 2005, YouTube has become the most successful Internet site providing a new generation of short video sharing service. Today, YouTube alone comprises approximately 20% of all HTTP traffic, or nearly 10% of all traffic on the…

Networking and Internet Architecture · Computer Science 2007-07-26 Xu Cheng , Cameron Dale , Jiangchuan Liu

YouTube (http://www.youtube.com) is an online, public-access video-sharing site that allows users to post short streaming-video submissions for open viewing. Along with Google, MySpace, Facebook, etc. it is one of the great success stories…

Physics Education · Physics 2008-08-27 A. P. Micolich

YouTube has become the second most popular website according to Alexa, and it represents an enticing platform for scammers to attract victims. Because of the computational difficulty of classifying multimedia, identifying scams on YouTube…

Cryptography and Security · Computer Science 2021-04-15 Elijah Bouma-Sims , Brad Reaves

We contribute the first publicly available dataset of factual claims from different platforms and fake YouTube videos on the 2023 Israel-Hamas war for automatic fake YouTube video classification. The FakeClaim data is collected from 60…

Information Retrieval · Computer Science 2024-01-31 Gautam Kishore Shahi , Amit Kumar Jaiswal , Thomas Mandl

Progress in AI is driven largely by the scale and quality of training data. Despite this, there is a deficit of empirical analysis examining the attributes of well-established datasets beyond text. In this work we conduct the largest and…

Dockerfiles are one of the most prevalent kinds of DevOps artifacts used in industry. Despite their prevalence, there is a lack of sophisticated semantics-aware static analysis of Dockerfiles. In this paper, we introduce a dataset of…

Software Engineering · Computer Science 2020-03-31 Jordan Henkel , Christian Bird , Shuvendu K. Lahiri , Thomas Reps

In this paper, we present a dataset of 713k articles collected between 02/2018-11/2018. These articles are collected directly from 194 news and media outlets including mainstream, hyper-partisan, and conspiracy sources. We incorporate…

Computers and Society · Computer Science 2019-04-03 Jeppe Norregaard , Benjamin D. Horne , Sibel Adali

Short video platforms, such as YouTube, Instagram, or TikTok, are used by billions of users. These platforms expose users to harmful content, ranging from clickbait or physical harms to hate or misinformation. Yet, we lack a comprehensive…

Computer Vision and Pattern Recognition · Computer Science 2025-04-24 Wonjeong Jo , Magdalena Wojcieszak

YouTube is the leading social media platform for sharing videos. As a result, it is plagued with misleading content that includes staged videos presented as real footages from an incident, videos with misrepresented context and videos where…

Computation and Language · Computer Science 2019-01-28 Priyank Palod , Ayush Patwari , Sudhanshu Bahety , Saurabh Bagchi , Pawan Goyal

Social information networks, such as YouTube, contains traces of both explicit online interaction (such as "like", leaving a comment, or subscribing to video feed), and latent interactions (such as quoting, or remixing parts of a video). We…

Social and Information Networks · Computer Science 2013-05-14 Lexing Xie , Apostol Natsev , Xuming He , John Kender , Matthew Hill , John R Smith

To facilitate the research on intelligent and human-like chatbots with multi-modal context, we introduce a new video-based multi-modal dialogue dataset, called TikTalk. We collect 38K videos from a popular video-sharing platform, along with…

Computation and Language · Computer Science 2023-09-11 Hongpeng Lin , Ludan Ruan , Wenke Xia , Peiyu Liu , Jingyuan Wen , Yixin Xu , Di Hu , Ruihua Song , Wayne Xin Zhao , Qin Jin , Zhiwu Lu

Here we present a massive longitudinal dataset of public Telegram content, comprising over 5.9 billion messages dating from 2015 to 2025, collected from 712 thousand channels and groups, enriched with metadata on forwards, reactions, and…