English
Related papers

Related papers: YFCC100M: The New Data in Multimedia Research

200 papers

Developing new ideas and algorithms in the fields of graph processing and relational learning requires public datasets. While Wikidata is the largest open source knowledge graph, involving more than fifty million entities, it is larger than…

Machine Learning · Computer Science 2019-10-07 Armand Boschin , Thomas Bonald

Film, a classic image style, is culturally significant to the whole photographic industry since it marks the birth of photography. However, film photography is time-consuming and expensive, necessitating a more efficient method for…

Computer Vision and Pattern Recognition · Computer Science 2023-11-06 Zinuo Li , Xuhang Chen , Shuqiang Wang , Chi-Man Pun

Interest in the research areas related to meme propagation and generation has been increasing rapidly in the last couple of years. Meme datasets available online are either specific to a context or contain no class information. Here, we…

Computation and Language · Computer Science 2019-12-04 Suryatej Reddy Vyalla , Vishaal Udandarao , Tanmoy Chakraborty

Traditional open-access datasets focusing on surgical procedures are often limited by their small size, typically consisting of fewer than 100 videos and less than 30 hours of footage, which leads to poor model generalization. To address…

Computer Vision and Pattern Recognition · Computer Science 2026-03-24 Chengan Che , Chao Wang , Tom Vercauteren , Sophia Tsoka , Luis C. Garcia-Peraza-Herrera

The rise of multi-million-item dataset initiatives has enabled data-hungry machine learning algorithms to reach near-human semantic classification at tasks such as object and scene recognition. Here we describe the Places Database, a…

Computer Vision and Pattern Recognition · Computer Science 2016-10-10 Bolei Zhou , Aditya Khosla , Agata Lapedriza , Antonio Torralba , Aude Oliva

The convergence of 3D geometric perception and video synthesis has created an unprecedented demand for large-scale video data that is rich in both semantic and spatio-temporal information. While existing datasets have advanced either 3D…

Computer Vision and Pattern Recognition · Computer Science 2026-04-28 Yunnan Wang , Kecheng Zheng , Jianyuan Wang , Minghao Chen , David Novotny , Christian Rupprecht , Yinghao Xu , Xing Zhu , Wenjun Zeng , Xin Jin , Yujun Shen

Current version identification (VI) datasets often lack sufficient size and musical diversity to train robust neural networks (NNs). Additionally, their non-representative clique size distributions prevent realistic system evaluations. To…

Sound · Computer Science 2024-10-24 R. Oguz Araz , Xavier Serra , Dmitry Bogdanov

Neural networks trained on datasets such as ImageNet have led to major advances in visual object classification. One obstacle that prevents networks from reasoning more deeply about complex scenes and situations, and from integrating visual…

The recent surge in open-source text-to-video generation models has significantly energized the research community, yet their dependence on proprietary training datasets remains a key constraint. While existing open datasets like Koala-36M…

Computer Vision and Pattern Recognition · Computer Science 2025-04-22 Xianpan Zhou

The arXiv has collected 1.5 million pre-print articles over 28 years, hosting literature from scientific fields including Physics, Mathematics, and Computer Science. Each pre-print features text, figures, authors, citations, categories, and…

Information Retrieval · Computer Science 2019-05-02 Colin B. Clement , Matthew Bierbaum , Kevin P. O'Keeffe , Alexander A. Alemi

We present NEWSROOM, a summarization dataset of 1.3 million articles and summaries written by authors and editors in newsrooms of 38 major news publications. Extracted from search and social media metadata between 1998 and 2017, these…

Computation and Language · Computer Science 2020-05-19 Max Grusky , Mor Naaman , Yoav Artzi

Multi-modal language-vision models trained on hundreds of millions of image-text pairs (e.g. CLIP, DALL-E) gained a recent surge, showing remarkable capability to perform zero- or few-shot learning and transfer even in absence of per-sample…

Computer Vision and Pattern Recognition · Computer Science 2021-11-04 Christoph Schuhmann , Richard Vencu , Romain Beaumont , Robert Kaczmarczyk , Clayton Mullis , Aarush Katta , Theo Coombes , Jenia Jitsev , Aran Komatsuzaki

Museums, libraries, and other cultural institutions continue to prioritize and build web-based visualization systems that increase access and discovery to digitized archives. Prominent examples exist that illustrate impressive…

Human-Computer Interaction · Computer Science 2020-09-07 Taylor Arnold , Nathaniel Ayers , Justin Madron , Robert Nelson , Lauren Tilton

We contribute the first large-scale dataset of scene sketches, SketchyScene, with the goal of advancing research on sketch understanding at both the object and scene level. The dataset is created through a novel and carefully designed…

Computer Vision and Pattern Recognition · Computer Science 2018-08-08 Changqing Zou , Qian Yu , Ruofei Du , Haoran Mo , Yi-Zhe Song , Tao Xiang , Chengying Gao , Baoquan Chen , Hao Zhang

The large abundance of perspective camera datasets facilitated the emergence of novel learning-based strategies for various tasks, such as camera localization, single image depth estimation, or view synthesis. However, panoramic or…

Computer Vision and Pattern Recognition · Computer Science 2024-07-08 Kibaek Park , Francois Rameau , Jaesik Park , In So Kweon

Recent years have witnessed a rapid development of immersive multimedia which bridges the gap between the real world and virtual space. Volumetric videos, as an emerging representative 3D video paradigm that empowers extended reality, stand…

Multimedia · Computer Science 2023-04-18 Kaiyuan Hu , Yili Jin , Haowen Yang , Junhua Liu , Fangxin Wang

In this paper we present a near-complete dataset of over 3M videos from 61K channels over 2.5 years (June 2019 to December 2021) from the social video hosting platform BitChute, a commonly used alternative to YouTube. Additionally, we…

Social and Information Networks · Computer Science 2022-02-14 Milo Trujillo , Maurício Gruppi , Cody Buntain , Benjamin D. Horne

Controlling the movements of dynamic objects and the camera within generated videos is a meaningful yet challenging task. Due to the lack of datasets with comprehensive 6D pose annotations, existing text-to-video methods can not…

Computer Vision and Pattern Recognition · Computer Science 2025-07-22 Xincheng Shuai , Henghui Ding , Zhenyuan Qin , Hao Luo , Xingjun Ma , Dacheng Tao

Multi-target multi-camera tracking is a crucial task that involves identifying and tracking individuals over time using video streams from multiple cameras. This task has practical applications in various fields, such as visual…

Computer Vision and Pattern Recognition · Computer Science 2024-04-01 Sanghyun Woo , Kwanyong Park , Inkyu Shin , Myungchul Kim , In So Kweon

Most existing large-scale academic search engines are built to retrieve text-based information. However, there are no large-scale retrieval services for scientific figures and tables. One challenge for such services is understanding…

Artificial Intelligence · Computer Science 2023-01-31 Zeba Karishma , Shaurya Rohatgi , Kavya Shrinivas Puranik , Jian Wu , C. Lee Giles