中文
相关论文

相关论文: Class Density and Dataset Quality in High-Dimensio…

200 篇论文

Quantifying the distance between datasets is a fundamental question in mathematics and machine learning. We propose \textit{magnitude distance}, a novel distance metric defined on finite datasets using the notion of the \emph{magnitude} of…

机器学习 · 计算机科学 2026-02-10 Sahel Torkamani , Henry Gouk , Rik Sarkar

Density Based Clustering are a type of Clustering methods using in data mining for extracting previously unknown patterns from data sets. There are a number of density based clustering methods such as DBSCAN, OPTICS, DENCLUE, VDBSCAN,…

机器学习 · 计算机科学 2023-07-25 Rupanka Bhuyan , Samarjeet Borah

Data-oriented applications, their users, and even the law require data of high quality. Research has divided the rather vague notion of data quality into various dimensions, such as accuracy, consistency, and reputation. To achieve the goal…

数据库 · 计算机科学 2024-12-09 Sedir Mohammed , Lisa Ehrlinger , Hazar Harmouch , Felix Naumann , Divesh Srivastava

Density-based Out-of-distribution (OOD) detection has recently been shown unreliable for the task of detecting OOD images. Various density ratio based approaches achieve good empirical performance, however methods typically lack a…

机器学习 · 统计学 2022-06-09 Mingtian Zhang , Andi Zhang , Tim Z. Xiao , Yitong Sun , Steven McDonagh

In this paper, we use a probabilistic model to estimate the number of uncorrelated features in a large dataset. Our model allows for both pairwise feature correlation (collinearity) and interdependency of multiple features…

机器学习 · 计算机科学 2023-09-26 Ghurumuruhan Ganesan

Real-world datasets are often of high dimension and effected by the curse of dimensionality. This hinders their comprehensibility and interpretability. To reduce the complexity feature selection aims to identify features that are crucial to…

机器学习 · 计算机科学 2023-04-18 Maximilian Stubbemann , Tobias Hille , Tom Hanika

We study compressible types in the context of (local and global) NIP. By extending a result in machine learning theory (the existence of a bound on the recursive teaching dimension), we prove density of compressible types. Using this, we…

逻辑 · 数学 2026-04-02 Martin Bays , Itay Kaplan , Pierre Simon

Classification and clustering are both important topics in statistical learning. A natural question herein is whether predefined classes are really different from one another, or whether clusters are really there. Specifically, we may be…

机器学习 · 统计学 2015-09-22 Qiyi Lu , Xingye Qiao

Data diversity is crucial for the instruction tuning of large language models. Existing studies have explored various diversity-aware data selection methods to construct high-quality datasets and enhance model performance. However, the…

计算与语言 · 计算机科学 2025-06-03 Yuming Yang , Yang Nan , Junjie Ye , Shihan Dou , Xiao Wang , Shuo Li , Huijie Lv , Mingqi Wu , Tao Gui , Qi Zhang , Xuanjing Huang

Complexity measures are essential to understand complex systems and there are numerous definitions to analyze one-dimensional data. However, extensions of these approaches to two or higher-dimensional data, such as images, are much less…

数据分析、统计与概率 · 物理学 2012-12-27 H. V. Ribeiro , L. Zunino , E. K. Lenzi , P. A. Santoro , R. S. Mendes

High dimensional data analysis is known to be as a challenging problem. In this article, we give a theoretical analysis of high dimensional classification of Gaussian data which relies on a geometrical analysis of the error measure. It…

统计理论 · 数学 2008-07-10 Robin Girard

This paper considers the problem of specifying a simple approximating density function for a given data set (x_1,...,x_n). Simplicity is measured by the number of modes but several different definitions of approximation are introduced. The…

统计理论 · 数学 2007-06-13 P. Laurie Davies , Arne Kovac

We aim to select data subsets for the fine-tuning of large language models to more effectively follow instructions. Prior work has emphasized the importance of diversity in dataset curation but relied on heuristics such as the number of…

机器学习 · 计算机科学 2024-02-07 Peiqi Wang , Yikang Shen , Zhen Guo , Matthew Stallone , Yoon Kim , Polina Golland , Rameswar Panda

We introduce the concept of a class of equivalence of molecular clouds represented by an abstract spherically symmetric, isotropic object. This object is described by use of abstract scales in respect to a given mass density distribution.…

星系天体物理 · 物理学 2017-01-18 Sava Donkov , Todor V. Veltchev , Ralf S. Klessen

Learning classifiers using skewed or imbalanced datasets can occasionally lead to classification issues; this is a serious issue. In some cases, one class contains the majority of examples while the other, which is frequently the more…

机器学习 · 计算机科学 2022-11-11 Satyendra Singh Rawat , Amit Kumar Mishra

Anomaly detection is not an easy problem since distribution of anomalous samples is unknown a priori. We explore a novel method that gives a trade-off possibility between one-class and two-class approaches, and leads to a better performance…

When the cost of misclassifying a sample is high, it is useful to have an accurate estimate of uncertainty in the prediction for that sample. There are also multiple types of uncertainty which are best estimated in different ways, for…

机器学习 · 计算机科学 2019-03-18 Richard Harang , Ethan M. Rudd

This paper studies the classification of high-dimensional Gaussian signals from low-dimensional noisy, linear measurements. In particular, it provides upper bounds (sufficient conditions) on the number of measurements required to drive the…

信息论 · 计算机科学 2016-11-03 Hugo Reboredo , Francesco Renna , Robert Calderbank , Miguel R. D. Rodrigues

Accurate approximation of a real-valued function depends on two aspects of the available data: the density of inputs within the domain of interest and the variation of the outputs over that domain. There are few methods for assessing…

数值分析 · 数学 2024-11-11 Andrew Gillette , Eugene Kur

We propose a new method to count objects of specific categories that are significantly smaller than the ground sampling distance of a satellite image. This task is hard due to the cluttered nature of scenes where different object categories…

计算机视觉与模式识别 · 计算机科学 2018-09-21 Andres C. Rodriguez , Jan D. Wegner