English
Related papers

Related papers: YouTube-8M: A Large-Scale Video Classification Ben…

200 papers

Accelerated by the tremendous increase in Internet bandwidth and storage space, video data has been generated, published and spread explosively, becoming an indispensable part of today's big data. In this paper, we focus on reviewing two…

Computer Vision and Pattern Recognition · Computer Science 2018-02-23 Zuxuan Wu , Ting Yao , Yanwei Fu , Yu-Gang Jiang

In recent years, the rapid expansion of dataset sizes and the increasing complexity of deep learning models have significantly escalated the demand for computational resources, both for data storage and model training. Dataset distillation…

Computer Vision and Pattern Recognition · Computer Science 2025-12-10 Zhe Li , Hadrien Reynaud , Mischa Dombrowski , Sarah Cechnicka , Franciskus Xaverius Erick , Bernhard Kainz

Understanding movies and their structural patterns is a crucial task in decoding the craft of video editing. While previous works have developed tools for general analysis, such as detecting characters or recognizing cinematography…

Computer Vision and Pattern Recognition · Computer Science 2022-10-25 Alejandro Pardo , Fabian Caba Heilbron , Juan León Alcázar , Ali Thabet , Bernard Ghanem

Owing to recent advances in machine learning and the ability to harvest large amounts of data during robotic-assisted surgeries, surgical data science is ripe for foundational work. We present a large dataset of surgical videos and their…

Computer Vision and Pattern Recognition · Computer Science 2026-05-11 Aneeq Zia , Max Berniker , Rogerio Nespolo , Xiaorui Zhang , Conor Perreault , Ziheng Wang , Benjamin Mueller , Ryan Schmidt , Kiran Bhattacharyya , Xi Liu , Anthony Jarc

The quality of video-text pairs fundamentally determines the upper bound of text-to-video models. Currently, the datasets used for training these models suffer from significant shortcomings, including low temporal consistency, poor-quality…

Computer Vision and Pattern Recognition · Computer Science 2024-08-06 Zhiyu Tan , Xiaomeng Yang , Luozheng Qin , Hao Li

YouTube is the leading social media platform for sharing videos. As a result, it is plagued with misleading content that includes staged videos presented as real footages from an incident, videos with misrepresented context and videos where…

Computation and Language · Computer Science 2019-01-28 Priyank Palod , Ayush Patwari , Sudhanshu Bahety , Saurabh Bagchi , Pawan Goyal

The potential for agents, whether embodied or software, to learn by observing other agents performing procedures involving objects and actions is rich. Current research on automatic procedure learning heavily relies on action labels or…

Computer Vision and Pattern Recognition · Computer Science 2017-11-23 Luowei Zhou , Chenliang Xu , Jason J. Corso

Humans share a strong tendency to memorize/forget some of the visual information they encounter. This paper focuses on providing computational models for the prediction of the intrinsic memorability of visual content. To address this new…

Computer Vision and Pattern Recognition · Computer Science 2024-02-28 Romain Cohendet , Claire-Hélène Demarty , Ngoc Q. K. Duong , Martin Engilberge

The availability of large-scale image captioning and visual question answering datasets has contributed significantly to recent successes in vision-and-language pre-training. However, these datasets are often collected with overrestrictive…

Computer Vision and Pattern Recognition · Computer Science 2021-03-31 Soravit Changpinyo , Piyush Sharma , Nan Ding , Radu Soricut

This paper introduces InternVid, a large-scale video-centric multimodal dataset that enables learning powerful and transferable video-text representations for multimodal understanding and generation. The InternVid dataset contains over 7…

Computer Vision and Pattern Recognition · Computer Science 2024-01-05 Yi Wang , Yinan He , Yizhuo Li , Kunchang Li , Jiashuo Yu , Xin Ma , Xinhao Li , Guo Chen , Xinyuan Chen , Yaohui Wang , Conghui He , Ping Luo , Ziwei Liu , Yali Wang , Limin Wang , Yu Qiao

Multi-modal models are data hungry. While datasets with natural images are abundant, medical image datasets can not afford the same luxury. To enable representation learning for medical images at scale, we turn to YouTube, a platform with a…

Computer Vision and Pattern Recognition · Computer Science 2026-01-15 Wisdom O. Ikezogwo , Kevin Zhang , Mehmet Saygin Seyfioglu , Fatemeh Ghezloo , Linda Shapiro , Ranjay Krishna

Data is the engine of modern computer vision, which necessitates collecting large-scale datasets. This is expensive, and guaranteeing the quality of the labels is a major challenge. In this paper, we investigate efficient annotation…

Computer Vision and Pattern Recognition · Computer Science 2021-04-27 Yuan-Hong Liao , Amlan Kar , Sanja Fidler

We introduce GLAMI-1M: the largest multilingual image-text classification dataset and benchmark. The dataset contains images of fashion products with item descriptions, each in 1 of 13 languages. Categorization into 191 classes has…

Computer Vision and Pattern Recognition · Computer Science 2022-11-29 Vaclav Kosar , Antonín Hoskovec , Milan Šulc , Radek Bartyzal

Learning-based visual data compression and analysis have attracted great interest from both academia and industry recently. More training as well as testing datasets, especially good quality video datasets are highly desirable for related…

Image and Video Processing · Electrical Eng. & Systems 2021-05-14 Xiaozhong Xu , Shan Liu , Zeqiang Li

Creating and labelling datasets of videos for use in training Human Activity Recognition models is an arduous task. In this paper, we approach this by using 3D rendering tools to generate a synthetic dataset of videos, and show that a…

Computer Vision and Pattern Recognition · Computer Science 2020-07-23 Ollie Matthews , Koki Ryu , Tarun Srivastava

Deep learning algorithms have pushed the boundaries of computer vision research and have depicted commendable performance in a variety of applications. However, training a robust deep neural network necessitates a large amount of labeled…

Computer Vision and Pattern Recognition · Computer Science 2023-07-13 Debanjan Goswami , Shayok Chakraborty

Being data-driven is one of the most iconic properties of deep learning algorithms. The birth of ImageNet drives a remarkable trend of "learning from large-scale data" in computer vision. Pretraining on ImageNet to obtain rich universal…

Computer Vision and Pattern Recognition · Computer Science 2023-03-13 Xianggang Yu , Mutian Xu , Yidan Zhang , Haolin Liu , Chongjie Ye , Yushuang Wu , Zizheng Yan , Chenming Zhu , Zhangyang Xiong , Tianyou Liang , Guanying Chen , Shuguang Cui , Xiaoguang Han

This work explores whether a deep generative model can learn complex knowledge solely from visual input, in contrast to the prevalent focus on text-based models like large language models (LLMs). We develop VideoWorld, an auto-regressive…

Computer Vision and Pattern Recognition · Computer Science 2025-03-06 Zhongwei Ren , Yunchao Wei , Xun Guo , Yao Zhao , Bingyi Kang , Jiashi Feng , Xiaojie Jin

Web-scale visual entity recognition, the task of associating images with their corresponding entities within vast knowledge bases like Wikipedia, presents significant challenges due to the lack of clean, large-scale training data. In this…

Computer Vision and Pattern Recognition · Computer Science 2024-11-01 Mathilde Caron , Alireza Fathi , Cordelia Schmid , Ahmet Iscen

A major challenge in text-video and text-audio retrieval is the lack of large-scale training data. This is unlike image-captioning, where datasets are in the order of millions of samples. To close this gap we propose a new video mining…

Computer Vision and Pattern Recognition · Computer Science 2022-04-05 Arsha Nagrani , Paul Hongsuck Seo , Bryan Seybold , Anja Hauth , Santiago Manen , Chen Sun , Cordelia Schmid
‹ Prev 1 3 4 5 6 7 10 Next ›