中文
相关论文

相关论文: The HASYv2 dataset

200 篇论文

We present the Noisy Ostracods, a noisy dataset for genus and species classification of crustacean ostracods with specialists' annotations. Over the 71466 specimens collected, 5.58% of them are estimated to be noisy (possibly problematic)…

机器学习 · 计算机科学 2024-12-04 Jiamian Hu , Yuanyuan Hong , Yihua Chen , He Wang , Moriaki Yasuhara

The MNIST dataset has become a standard benchmark for learning, classification and computer vision systems. Contributing to its widespread adoption are the understandable and intuitive nature of the task, its relatively small size and…

计算机视觉与模式识别 · 计算机科学 2017-03-02 Gregory Cohen , Saeed Afshar , Jonathan Tapson , André van Schaik

Semi-iNat is a challenging dataset for semi-supervised classification with a long-tailed distribution of classes, fine-grained categories, and domain shifts between labeled and unlabeled data. This dataset is behind the second iteration of…

计算机视觉与模式识别 · 计算机科学 2021-06-24 Jong-Chyi Su , Subhransu Maji

In this work, we introduce a practical dataset named HUST bearing, that provides a large set of vibration data on different ball bearings. This dataset contains 90 raw vibration data of 6 types of defects (inner crack, outer crack, ball…

机器学习 · 计算机科学 2023-10-03 Nguyen Duc Thuan , Hoang Si Hong

Handwritten document image binarization is challenging due to high variability in the written content and complex background attributes such as page style, paper quality, stains, shadow gradients, and non-uniform illumination. While the…

计算机视觉与模式识别 · 计算机科学 2021-11-04 Kaustubh Sadekar , Ashish Tiwari , Prajwal Singh , Shanmuganathan Raman

In our study, we conducted a comprehensive analysis of three widely used datasets in the domain of building footprint extraction using deep neural networks: the INRIA Aerial Image Labelling dataset, SpaceNet 2: Building Detection v2, and…

计算机视觉与模式识别 · 计算机科学 2026-04-22 Yeshwanth Kumar Adimoolam , Charalambos Poullis , Melinos Averkiou

We present Open Images V4, a dataset of 9.2M images with unified annotations for image classification, object detection and visual relationship detection. The images have a Creative Commons Attribution license that allows to share and adapt…

In this work, we announce a comprehensive well curated and opensource dataset with millions of samples for pre-college and college level problems in mathematicsand science. A preliminary set of results using transformer architecture with…

数值分析 · 数学 2021-10-01 Neeraj Kollepara , Snehith Kumar Chatakonda , Pawan Kumar

This document describes the details and the motivation behind a new dataset we collected for the semi-supervised recognition challenge~\cite{semi-aves} at the FGVC7 workshop at CVPR 2020. The dataset contains 1000 species of birds sampled…

计算机视觉与模式识别 · 计算机科学 2021-03-15 Jong-Chyi Su , Subhransu Maji

In this paper, we present MusPy, an open source Python library for symbolic music generation. MusPy provides easy-to-use tools for essential components in a music generation system, including dataset management, data I/O, data preprocessing…

声音 · 计算机科学 2020-08-06 Hao-Wen Dong , Ke Chen , Julian McAuley , Taylor Berg-Kirkpatrick

The MNIST dataset containing thousands of handwritten digit images is still a fundamental benchmark for evaluating various pattern-recognition and image-classification models. Linear separability is a key concept in many statistical and…

机器学习 · 计算机科学 2026-03-16 Ákos Hajnal

We present an interesting and challenging dataset that features a large number of scenes with messy tables captured from multiple camera views. Each scene in this dataset is highly complex, containing multiple object instances that could be…

计算机视觉与模式识别 · 计算机科学 2020-07-30 Zhongang Cai , Junzhe Zhang , Daxuan Ren , Cunjun Yu , Haiyu Zhao , Shuai Yi , Chai Kiat Yeo , Chen Change Loy

Unit testing is an essential part of the software development process, which helps to identify issues with source code in early stages of development and prevent regressions. Machine learning has emerged as viable approach to help software…

软件工程 · 计算机科学 2022-03-25 Michele Tufano , Shao Kun Deng , Neel Sundaresan , Alexey Svyatkovskiy

This study introduces a federated learning-based approach to predict HER2 status from hematoxylin and eosin (HE)-stained whole slide images (WSIs), reducing costs and speeding up treatment decisions. To address label imbalance and feature…

计算机视觉与模式识别 · 计算机科学 2024-12-20 Kamorudeen A. Amuda , Almustapha A. Wakili

We introduce a large-scale dataset of the complete texts of free/open source software (FOSS) license variants. To assemble it we have collected from the Software Heritage archive-the largest publicly available archive of FOSS source code…

软件工程 · 计算机科学 2022-04-04 Stefano Zacchiroli

We present Typography-MNIST (TMNIST), a dataset comprising of 565,292 MNIST-style grayscale images representing 1,812 unique glyphs in varied styles of 1,355 Google-fonts. The glyph-list contains common characters from over 150 of the…

计算机视觉与模式识别 · 计算机科学 2022-02-17 Nimish Magre , Nicholas Brown

Recognising animals based on distinctive body patterns, such as stripes, spots, or other markings, in night images is a complex task in computer vision. Existing methods for detecting animals in images often rely on colour information,…

计算机视觉与模式识别 · 计算机科学 2024-10-29 John Atanbori

Images in visualization publications contain rich information, e.g., novel visualization designs and implicit design patterns of visualizations. A systematic collection of these images can contribute to the community in many aspects, such…

计算机视觉与模式识别 · 计算机科学 2022-03-08 Dazhen Deng , Yihong Wu , Xinhuan Shu , Jiang Wu , Siwei Fu , Weiwei Cui , Yingcai Wu

Existing image classification datasets used in computer vision tend to have a uniform distribution of images across object categories. In contrast, the natural world is heavily imbalanced, as some species are more abundant and easier to…

计算机视觉与模式识别 · 计算机科学 2018-04-12 Grant Van Horn , Oisin Mac Aodha , Yang Song , Yin Cui , Chen Sun , Alex Shepard , Hartwig Adam , Pietro Perona , Serge Belongie

In this paper, we present ManyTypes4Py, a large Python dataset for machine learning (ML)-based type inference. The dataset contains a total of 5,382 Python projects with more than 869K type annotations. Duplicate source code files were…

软件工程 · 计算机科学 2021-04-13 Amir M. Mir , Evaldas Latoskinas , Georgios Gousios