中文
相关论文

相关论文: Class Density and Dataset Quality in High-Dimensio…

200 篇论文

Model selection consistency in the high-dimensional regression setting can be achieved only if strong assumptions are fulfilled. We therefore suggest to pursue a different goal, which we call a minimal class of models. The minimal class of…

统计方法学 · 统计学 2015-11-26 Daniel Nevo , Ya'acov Ritov

Dimensionality reduction is a crucial technique in data analysis, as it allows for the efficient visualization and understanding of high-dimensional datasets. The circular coordinate is one of the topological data analysis techniques…

代数拓扑 · 数学 2023-01-31 Taejin Paik , Jaemin Park

The random coefficients model is an extension of the linear regression model that allows for unobserved heterogeneity in the population by modeling the regression coefficients as random variables. Given data from this model, the statistical…

统计方法学 · 统计学 2018-03-15 Fabian Dunker , Konstantin Eckle , Katharina Proksch , Johannes Schmidt-Hieber

In this paper, we investigate the effect of addressing difficult samples from a given text dataset on the downstream text classification task. We define difficult samples as being non-obvious cases for text classification by analysing them…

计算与语言 · 计算机科学 2023-02-14 Shashank Mujumdar , Stuti Mehta , Hima Patel , Suman Mitra

High density clusters can be characterized by the connected components of a level set $L(\lambda) = \{x:\ p(x)>\lambda\}$ of the underlying probability density function $p$ generating the data, at some appropriate level $\lambda\geq 0$. The…

机器学习 · 统计学 2010-11-15 Alessandro Rinaldo , Aarti Singh , Rebecca Nugent , Larry Wasserman

Breast density classification is an essential part of breast cancer screening. Although a lot of prior work considered this problem as a task for learning algorithms, to our knowledge, all of them used small and not clinically realistic…

计算机视觉与模式识别 · 计算机科学 2017-11-13 Nan Wu , Krzysztof J. Geras , Yiqiu Shen , Jingyi Su , S. Gene Kim , Eric Kim , Stacey Wolfson , Linda Moy , Kyunghyun Cho

Data Science and Machine Learning have become fundamental assets for companies and research institutions alike. As one of its fields, supervised classification allows for class prediction of new samples, learning from given training data.…

Many current applications in data science need rich model classes to adequately represent the statistics that may be driving the observations. But rich model classes may be too complex to admit estimators that converge to the truth with…

信息论 · 计算机科学 2022-05-04 N. Santhanam , V. Anantharam , W. Szpankowski

The quality of data is context dependent. Starting from this intuition and experience, we propose and develop a conceptual framework that captures in formal terms the notion of "context-dependent data quality". We start by proposing a…

数据库 · 计算机科学 2016-08-16 Leopoldo Bertossi , Flavio Rizzolo

As technology advanced, collecting data via automatic collection devices become popular, thus we commonly face data sets with lengthy variables, especially when these data sets are collected without specific research goals beforehand. It…

机器学习 · 统计学 2022-05-10 Wan-Ping Nicole Chen , Yuan-chin Ivan Chang

Clustering is an underspecified task: there are no universal criteria for what makes a good clustering. This is especially true for relational data, where similarity can be based on the features of individuals, the relationships between…

机器学习 · 统计学 2017-09-29 Sebastijan Dumancic , Hendrik Blockeel

Class-wise characteristics of training examples affect the performance of deep classifiers. A well-studied example is when the number of training examples of classes follows a long-tailed distribution, a situation that is likely to yield…

机器学习 · 计算机科学 2025-05-02 Z. S. Baltaci , K. Oksuz , S. Kuzucu , K. Tezoren , B. K. Konar , A. Ozkan , E. Akbas , S. Kalkan

Density estimation is an interdisciplinary topic at the intersection of statistics, theoretical computer science and machine learning. We review some old and new techniques for bounding the sample complexity of estimating densities of…

统计理论 · 数学 2018-02-23 Hassan Ashtiani , Abbas Mehrabian

Modern computer vision foundation models are trained on massive amounts of data, incurring large economic and environmental costs. Recent research has suggested that improving data quality can significantly reduce the need for data…

计算机视觉与模式识别 · 计算机科学 2023-11-08 Benjamin Feuer , Chinmay Hegde

A number of classification problems need to deal with data imbalance between classes. Often it is desired to have a high recall on the minority class while maintaining a high precision on the majority class. In this paper, we review a…

应用统计 · 统计学 2016-08-23 Ajinkya More

The estimation of probability densities based on available data is a central task in many statistical applications. Especially in the case of large ensembles with many samples or high-dimensional sample spaces, computationally efficient…

统计方法学 · 统计学 2017-05-04 Daniel W. Meyer

This brief paper develops a probability density that models processes for which the physical mechanism is unknown. It has desirable properties which are not realized by densities derived from Gaussian process or other classic methods. In…

综合物理 · 物理学 2011-04-21 Steven C. Gustafson , Adam C. Hillier

This paper considers estimation of a univariate density from an individual numerical sequence. It is assumed that (i) the limiting relative frequencies of the numerical sequence are governed by an unknown density, and (ii) there is a known…

概率论 · 数学 2008-06-19 Andrew B. Nobel , Gusztav Morvai , Sanjeev R. Kulkarni

The estimation of a density profile from experimental data points is a challenging problem, usually tackled by plotting a histogram. Prior assumptions on the nature of the density, from its smoothness to the specification of its form, allow…

统计方法学 · 统计学 2015-03-13 Alberto Bernacchia , Simone Pigolotti

When selecting data for training large-scale models, standard practice is to filter for examples that match human notions of data quality. Such filtering yields qualitatively clean datapoints that intuitively should improve model behavior.…

机器学习 · 计算机科学 2024-01-24 Logan Engstrom , Axel Feldmann , Aleksander Madry