English
Related papers

Related papers: Class Density and Dataset Quality in High-Dimensio…

200 papers

Quantifying the distance between datasets is a fundamental question in mathematics and machine learning. We propose \textit{magnitude distance}, a novel distance metric defined on finite datasets using the notion of the \emph{magnitude} of…

Machine Learning · Computer Science 2026-02-10 Sahel Torkamani , Henry Gouk , Rik Sarkar

Density Based Clustering are a type of Clustering methods using in data mining for extracting previously unknown patterns from data sets. There are a number of density based clustering methods such as DBSCAN, OPTICS, DENCLUE, VDBSCAN,…

Machine Learning · Computer Science 2023-07-25 Rupanka Bhuyan , Samarjeet Borah

Data-oriented applications, their users, and even the law require data of high quality. Research has divided the rather vague notion of data quality into various dimensions, such as accuracy, consistency, and reputation. To achieve the goal…

Databases · Computer Science 2024-12-09 Sedir Mohammed , Lisa Ehrlinger , Hazar Harmouch , Felix Naumann , Divesh Srivastava

Density-based Out-of-distribution (OOD) detection has recently been shown unreliable for the task of detecting OOD images. Various density ratio based approaches achieve good empirical performance, however methods typically lack a…

Machine Learning · Statistics 2022-06-09 Mingtian Zhang , Andi Zhang , Tim Z. Xiao , Yitong Sun , Steven McDonagh

In this paper, we use a probabilistic model to estimate the number of uncorrelated features in a large dataset. Our model allows for both pairwise feature correlation (collinearity) and interdependency of multiple features…

Machine Learning · Computer Science 2023-09-26 Ghurumuruhan Ganesan

Real-world datasets are often of high dimension and effected by the curse of dimensionality. This hinders their comprehensibility and interpretability. To reduce the complexity feature selection aims to identify features that are crucial to…

Machine Learning · Computer Science 2023-04-18 Maximilian Stubbemann , Tobias Hille , Tom Hanika

We study compressible types in the context of (local and global) NIP. By extending a result in machine learning theory (the existence of a bound on the recursive teaching dimension), we prove density of compressible types. Using this, we…

Logic · Mathematics 2026-04-02 Martin Bays , Itay Kaplan , Pierre Simon

Classification and clustering are both important topics in statistical learning. A natural question herein is whether predefined classes are really different from one another, or whether clusters are really there. Specifically, we may be…

Machine Learning · Statistics 2015-09-22 Qiyi Lu , Xingye Qiao

Data diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the…

Computation and Language · Computer Science 2025-06-03 Yuming Yang , Yang Nan , Junjie Ye , Shihan Dou , Xiao Wang , Shuo Li , Huijie Lv , Mingqi Wu , Tao Gui , Qi Zhang , Xuanjing Huang

Complexity measures are essential to understand complex systems and there are numerous definitions to analyze one-dimensional data. However, extensions of these approaches to two or higher-dimensional data, such as images, are much less…

Data Analysis, Statistics and Probability · Physics 2012-12-27 H. V. Ribeiro , L. Zunino , E. K. Lenzi , P. A. Santoro , R. S. Mendes

High dimensional data analysis is known to be as a challenging problem. In this article, we give a theoretical analysis of high dimensional classification of Gaussian data which relies on a geometrical analysis of the error measure. It…

Statistics Theory · Mathematics 2008-07-10 Robin Girard

This paper considers the problem of specifying a simple approximating density function for a given data set (x_1,...,x_n). Simplicity is measured by the number of modes but several different definitions of approximation are introduced. The…

Statistics Theory · Mathematics 2007-06-13 P. Laurie Davies , Arne Kovac

We aim to select data subsets for the fine-tuning of large language models to more effectively follow instructions. Prior work has emphasized the importance of diversity in dataset curation but relied on heuristics such as the number of…

Machine Learning · Computer Science 2024-02-07 Peiqi Wang , Yikang Shen , Zhen Guo , Matthew Stallone , Yoon Kim , Polina Golland , Rameswar Panda

We introduce the concept of a class of equivalence of molecular clouds represented by an abstract spherically symmetric, isotropic object. This object is described by use of abstract scales in respect to a given mass density distribution.…

Astrophysics of Galaxies · Physics 2017-01-18 Sava Donkov , Todor V. Veltchev , Ralf S. Klessen

Learning classifiers using skewed or imbalanced datasets can occasionally lead to classification issues; this is a serious issue. In some cases, one class contains the majority of examples while the other, which is frequently the more…

Machine Learning · Computer Science 2022-11-11 Satyendra Singh Rawat , Amit Kumar Mishra

Anomaly detection is not an easy problem since distribution of anomalous samples is unknown a priori. We explore a novel method that gives a trade-off possibility between one-class and two-class approaches, and leads to a better performance…

Machine Learning · Statistics 2020-05-26 Maxim Borisyak , Artem Ryzhikov , Andrey Ustyuzhanin , Denis Derkach , Fedor Ratnikov , Olga Mineeva

When the cost of misclassifying a sample is high, it is useful to have an accurate estimate of uncertainty in the prediction for that sample. There are also multiple types of uncertainty which are best estimated in different ways, for…

Machine Learning · Computer Science 2019-03-18 Richard Harang , Ethan M. Rudd

This paper studies the classification of high-dimensional Gaussian signals from low-dimensional noisy, linear measurements. In particular, it provides upper bounds (sufficient conditions) on the number of measurements required to drive the…

Information Theory · Computer Science 2016-11-03 Hugo Reboredo , Francesco Renna , Robert Calderbank , Miguel R. D. Rodrigues

Accurate approximation of a real-valued function depends on two aspects of the available data: the density of inputs within the domain of interest and the variation of the outputs over that domain. There are few methods for assessing…

Numerical Analysis · Mathematics 2024-11-11 Andrew Gillette , Eugene Kur

We propose a new method to count objects of specific categories that are significantly smaller than the ground sampling distance of a satellite image. This task is hard due to the cluttered nature of scenes where different object categories…

Computer Vision and Pattern Recognition · Computer Science 2018-09-21 Andres C. Rodriguez , Jan D. Wegner
‹ Prev 1 4 5 6 7 8 10 Next ›