English
Related papers

Related papers: The HASYv2 dataset

200 papers

Software is used in critical applications in our day-to-day life and it is important to ensure its correctness. One popular approach to assess correctness is to evaluate software on tests. If a test fails, it indicates a fault in the…

Software Engineering · Computer Science 2025-04-01 Max Hort , Leon Moonen

We introduce NSF-SciFy, a comprehensive dataset of scientific claims and investigation proposals extracted from National Science Foundation award abstracts. While previous scientific claim verification datasets have been limited in size and…

Computation and Language · Computer Science 2026-05-27 Delip Rao , Weiqiu You , Eric Wong , Chris Callison-Burch

Quantifying the similarity between datasets has widespread applications in statistics and machine learning. The performance of a predictive model on novel datasets, referred to as generalizability, depends on how similar the training and…

Methodology · Statistics 2025-06-18 Marieke Stolte , Franziska Kappenberg , Jörg Rahnenführer , Andrea Bommert

We investigate the concept of deep barcodes and propose two methods to generate them in order to expedite the process of classification and retrieval of histopathology images. Since binary search is computationally less expensive, in terms…

Image and Video Processing · Electrical Eng. & Systems 2018-05-24 Meghana Dinesh Kumar , Morteza Babaie , Hamid Tizhoosh

A two-sample hypothesis test is a statistical procedure used to determine whether the distributions generating two samples are identical. We consider the two-sample testing problem in a new scenario where the sample measurements (or sample…

Machine Learning · Computer Science 2024-07-01 Weizhi Li , Prad Kadambi , Pouria Saidi , Karthikeyan Natesan Ramamurthy , Gautam Dasarathy , Visar Berisha

Generalizable manipulation skills, which can be composed to tackle long-horizon and complex daily chores, are one of the cornerstones of Embodied AI. However, existing benchmarks, mostly composed of a suite of simulatable environments, are…

We present OpenStaxQA, an evaluation benchmark specific to college-level educational applications based on 43 open-source college textbooks in English, Spanish, and Polish, available under a permissive Creative Commons license. We finetune…

Computation and Language · Computer Science 2025-10-09 Pranav Gupta

This paper presents an open-source dataset RflyMAD, a Multicopter Abnomal Dataset developed by Reliable Flight Control (Rfly) Group aiming to promote the development of research fields like fault detection and isolation (FDI) or health…

Robotics · Computer Science 2024-01-12 Xiangli Le , Bo Jin , Gen Cui , Xunhua Dai , Quan Quan

Binary classification is a task that involves the classification of data into one of two distinct classes. It is widely utilized in various fields. However, conventional classifiers tend to make overconfident predictions for data that…

Machine Learning · Computer Science 2025-03-13 Shoma Yokura , Akihisa Ichiki

The growing importance of data visualization in business intelligence and data science emphasizes the need for tools that can efficiently generate meaningful visualizations from large datasets. Existing tools fall into two main categories:…

Databases · Computer Science 2024-09-10 Yupeng Xie , Yuyu Luo , Guoliang Li , Nan Tang

This is the second in a series of three papers in which we present an end-to-end simulation from the MICE collaboration, the MICE Grand Challenge (MICE-GC) run. The N-body contains about 70 billion dark-matter particles in a $(3 \, h^{-1}…

Cosmology and Nongalactic Astrophysics · Physics 2015-10-26 M. Crocce , F. J. Castander , E. Gaztanaga , P. Fosalba , J. Carretero

Various large language models (LLMs) have been proposed in recent years, including closed- and open-source ones, continually setting new records on multiple benchmarks. However, the development of LLMs still faces several issues, such as…

Computation and Language · Computer Science 2024-04-05 Ruyi Gan , Ziwei Wu , Renliang Sun , Junyu Lu , Xiaojun Wu , Dixiang Zhang , Kunhao Pan , Junqing He , Yuanhe Tian , Ping Yang , Qi Yang , Hao Wang , Jiaxing Zhang , Yan Song

Tabular data synthesis is a long-standing research topic in machine learning. Many different methods have been proposed over the past decades, ranging from statistical methods to deep generative methods. However, it has not always been…

Machine Learning · Computer Science 2023-05-30 Jayoung Kim , Chaejeong Lee , Noseong Park

Hashing produces compact representations for documents, to perform tasks like classification or retrieval based on these short codes. When hashing is supervised, the codes are trained using labels on the training data. This paper first…

Computer Vision and Pattern Recognition · Computer Science 2017-08-11 Alexandre Sablayrolles , Matthijs Douze , Hervé Jégou , Nicolas Usunier

While large language models provide significant convenience for software development, they can lead to ethical issues in job interviews and student assignments. Therefore, determining whether a piece of code is written by a human or…

Software Engineering · Computer Science 2025-05-27 Basak Demirok , Mucahid Kutlu

With the rapid development of deep learning technology and improvement in computing capability, deep learning has been widely used in the field of hyperspectral image (HSI) classification. In general, deep learning models often contain many…

Computer Vision and Pattern Recognition · Computer Science 2021-12-06 Sen Jia , Shuguo Jiang , Zhijie Lin , Nanying Li , Meng Xu , Shiqi Yu

Many classification problems can be difficult to formulate directly in terms of the traditional supervised setting, where both training and test samples are individual feature vectors. There are cases in which samples are better described…

Machine Learning · Statistics 2016-07-12 Veronika Cheplygina , David M. J. Tax , Marco Loog

The ability of learning from noisy labels is very useful in many visual recognition tasks, as a vast amount of data with noisy labels are relatively easy to obtain. Traditionally, the label noises have been treated as statistical outliers,…

Computer Vision and Pattern Recognition · Computer Science 2017-04-11 Yuncheng Li , Jianchao Yang , Yale Song , Liangliang Cao , Jiebo Luo , Li-Jia Li

A key requirement for supervised machine learning is labeled training data, which is created by annotating unlabeled data with the appropriate class. Because this process can in many cases not be done by machines, labeling needs to be…

Machine Learning · Computer Science 2019-12-12 Nicolas Michael Müller , Karla Markert

Traffic signs are essential map features globally in the era of autonomous driving and smart cities. To develop accurate and robust algorithms for traffic sign detection and classification, a large-scale and diverse benchmark dataset is…

Computer Vision and Pattern Recognition · Computer Science 2020-05-08 Christian Ertler , Jerneja Mislej , Tobias Ollmann , Lorenzo Porzi , Gerhard Neuhold , Yubin Kuang
‹ Prev 1 8 9 10 Next ›