中文
相关论文

相关论文: Class Density and Dataset Quality in High-Dimensio…

200 篇论文

Methods of performing anomaly detection on high-dimensional data sets are needed, since algorithms which are trained on data are only expected to perform well on data that is similar to the training data. There are theoretical results on…

机器学习 · 计算机科学 2020-11-13 Forrest Laine , Claire Tomlin

In this paper, we introduce a new measure called Term_Class relevance to compute the relevancy of a term in classifying a document into a particular class. The proposed measure estimates the degree of relevance of a given term, in placing…

信息检索 · 计算机科学 2016-09-15 D S Guru , Mahamad Suhil

In machine learning, the performance of a classifier depends on both the classifier model and the dataset. For a specific neural network classifier, the training process varies with the training set used; some training data make training…

机器学习 · 计算机科学 2020-06-01 Shuyue Guan , Murray Loew , Hanseok Ko

We study finitely additive extensions of the asymptotic density to all the subsets of natural numbers. Such measures are called density measures. We consider a class of density measures constructed from free ultrafilters on $\mathbb{N}$ and…

数论 · 数学 2016-01-26 Ryoichi Kunisada

It has long been noticed that high dimension data exhibits strange patterns. This has been variously interpreted as either a "blessing" or a "curse", causing uncomfortable inconsistencies in the literature. We propose that these patterns…

计算机视觉与模式识别 · 计算机科学 2020-03-18 Wen-Yan Lin

Defect prediction is crucial for software quality assurance and has been extensively researched over recent decades. However, prior studies rarely focus on data complexity in defect prediction tasks, and even less on understanding the…

软件工程 · 计算机科学 2023-05-08 Xiaohui Wan , Zheng Zheng , Fangyun Qin , Xuhui Lu

It is now practically the norm for data to be very high dimensional in areas such as genetics, machine vision, image analysis and many others. When analyzing such data, parametric models are often too inflexible while nonparametric…

统计方法学 · 统计学 2011-05-31 Abhishek Bhattacharya , Garritt Page , David Dunson

We use a formal correspondence between thermodynamics and inference, where the number of samples can be thought of as the inverse temperature, to study a quantity called ``learning capacity'' which is a measure of the effective…

机器学习 · 计算机科学 2024-10-22 Daiwei Chen , Wei-Kai Chang , Pratik Chaudhari

Density-functional theory is a formally exact description of a many-body quantum system in terms of its density; in practice, however, approximations to the universal density functional are required. In this work, a model based on deep…

计算物理 · 物理学 2016-08-02 Jeffrey M. McMahon

Deep neural networks (DNNs) have achieved exceptional performances in many tasks, particularly, in supervised classification tasks. However, achievements with supervised classification tasks are based on large datasets with well-separated…

计算机视觉与模式识别 · 计算机科学 2018-02-06 Kazuma Arino , Yohei Kikuta

Quantile classifiers for potentially high-dimensional data are defined by classifying an observation according to a sum of appropriately weighted component-wise distances of the components of the observation to the within-class quantiles.…

统计方法学 · 统计学 2013-11-13 Christian Hennig , Cinzia Viroli

A new clustering accuracy measure is proposed to determine the unknown number of clusters and to assess the quality of clustering of a data set given in any dimensional space. Our validity index applies the classical nonparametric…

统计方法学 · 统计学 2022-02-15 Soumita Modak

Integrating datasets from different disciplines is hard because the data are often qualitatively different in meaning, scale, and reliability. When two datasets describe the same entities, many scientific questions can be phrased around…

This paper discusses an approach with machine-learning probability models to evaluate the difference between good and bad data quality in a dataset. A decision tree algorithm is used to predict data quality based on no domain knowledge of…

机器学习 · 计算机科学 2020-09-16 Allen ONeill

Classification tasks are usually analysed and improved through new model architectures or hyperparameter optimisation but the underlying properties of datasets are discovered on an ad-hoc basis as errors occur. However, understanding the…

计算与语言 · 计算机科学 2018-12-10 Edward Collins , Nikolai Rozanov , Bingbing Zhang

One of the fundamental problems in machine learning is the estimation of a probability distribution from data. Many techniques have been proposed to study the structure of data, most often building around the assumption that observations…

机器学习 · 统计学 2013-02-22 Oren Rippel , Ryan Prescott Adams

We propose reinterpreting copula density estimation as a discriminative task. Under this novel estimation scheme, we train a classifier to distinguish samples from the joint density from those of the product of independent marginals,…

统计方法学 · 统计学 2025-03-20 David Huk , Mark Steel , Ritabrata Dutta

Having a sufficient quantity of quality data is a critical enabler of training effective machine learning models. Being able to effectively determine the adequacy of a dataset prior to training and evaluating a model's performance would be…

机器学习 · 计算机科学 2026-04-28 Arya Hatamian , Lionel Levine , Haniyeh Ehsani Oskouie , Majid Sarrafzadeh

Deep Neural Networks (DNNs), with its promising performance, are being increasingly used in safety critical applications such as autonomous driving, cancer detection, and secure authentication. With growing importance in deep learning,…

机器学习 · 计算机科学 2019-11-19 Senthil Mani , Anush Sankaran , Srikanth Tamilselvam , Akshay Sethi

Classifying samples in incomplete datasets is a common aim for machine learning practitioners, but is non-trivial. Missing data is found in most real-world datasets and these missing values are typically imputed using established methods,…