中文
相关论文

相关论文: A comparison of Gap statistic definitions with and…

200 篇论文

Statistical depth is the act of gauging how representative a point is compared to a reference probability measure. The depth allows introducing rankings and orderings to data living in multivariate, or function spaces. Though widely applied…

统计理论 · 数学 2021-05-28 George Wynne , Stanislav Nagy

Comparison of three kind of the clustering and find cost function and loss function and calculate them. Error rate of the clustering methods and how to calculate the error percentage always be one on the important factor for evaluating the…

机器学习 · 计算机科学 2014-11-14 Kamran Kowsari

The Thue--Morse sequence is a prototypical automatic sequence found in diverse areas of mathematics, and in computer science. We study occurrences of factors $w$ within this sequence, more precisely, the sequence of gaps between consecutive…

组合数学 · 数学 2021-11-19 Lukas Spiegelhofer

In a standard cluster analysis, such as k-means, in addition to clusters locations and distances between them, it's important to know if they are connected or well separated from each other. The main focus of this paper is discovering the…

机器学习 · 统计学 2017-05-22 Evgeny Bauman , Konstantin Bauman

Spectral Clustering(SC) is a prominent data clustering technique of recent times which has attracted much attention from researchers. It is a highly data-driven method and makes no strict assumptions on the structure of the data to be…

机器学习 · 计算机科学 2019-09-18 Lalith Srikanth Chintalapati , Raghunatha Sarma Rachakonda

A hypothesis testing and an interval estimation are studied for the common mean of several lognormal populations. Two methods are given based on the concept of generalized p-value and generalized confidence interval. These new methods are…

统计理论 · 数学 2014-05-06 Javad Behboodian , Ali Akbar Jafari

A simple model to study subspace clustering is the high-dimensional $k$-Gaussian mixture model where the cluster means are sparse vectors. Here we provide an exact asymptotic characterization of the statistically optimal reconstruction…

机器学习 · 统计学 2023-04-04 Luca Pesce , Bruno Loureiro , Florent Krzakala , Lenka Zdeborová

In machine learning and data mining, Cluster analysis is one of the most widely used unsupervised learning technique. Philosophy of this algorithm is to find similar data items and group them together based on any distance function in…

机器学习 · 统计学 2018-10-09 Kumarjit Pathak , Jitin Kapila

For the decomposability property is very a practical one in Welfare analysis, most researchers and users favor decomposable poverty indices such as the Foster-Greer-Thorbeck poverty index. This may lead to neglect the so important weighted…

统计方法学 · 统计学 2012-01-12 Mohamed Cheikh Haidara , Gane Samb Lo

Conformal prediction is a popular technique for constructing prediction intervals with distribution-free coverage guarantees. The coverage is marginal, meaning it only holds on average over the entire population but not necessarily for any…

统计方法学 · 统计学 2026-05-28 Yao Zhang , Emmanuel J. Candès

Imagine that you could calculate of posttest probabilities, i.e. Bayes theorem with simple addition. This is possible if we stop thinking of probabilities as ranging from 0 to 1.0. There is a naturally occurring linear probability space…

其他统计学 · 统计学 2019-04-03 Christopher M Rembold

The classical paradigm of scoring rules is to discriminate between two different forecasts by comparing them with observations. The probability distribution of the observed record is assumed to be perfect as a verification benchmark. In…

统计方法学 · 统计学 2021-08-06 Julie Bessac , Philippe Naveau

The coefficient of variation is a useful indicator for comparing the spread of values between dataset with different units or widely different means. In this paper we address the problem of investigating the equality of the coefficients of…

统计方法学 · 统计学 2023-06-06 Francesco Bertolino , Silvia Columbu , Mara Manca , Monica Musio

The problem of clustering is considered, for the case when each data point is a sample generated by a stationary ergodic process. We propose a very natural asymptotic notion of consistency, and show that simple consistent algorithms exist,…

机器学习 · 计算机科学 2013-05-01 Daniil Ryabko

The problem of clustering is considered, for the case when each data point is a sample generated by a stationary ergodic process. We propose a very natural asymptotic notion of consistency, and show that simple consistent algorithms exist,…

机器学习 · 计算机科学 2010-05-31 Daniil Ryabko

The $\Delta_3(L)$ statistic of Random Matrix Theory is defined as the average of a set of random numbers $\{\delta\}$, derived from a spectrum. The distribution $p(\delta)$ of these random numbers is used as the basis of a maximum…

核理论 · 物理学 2015-03-19 Declan Mulhall

Clustering is a widely used unsupervised learning method for finding structure in the data. However, the resulting clusters are typically presented without any guarantees on their robustness; slightly changing the used data sample or…

机器学习 · 统计学 2017-01-02 Andreas Henelius , Kai Puolamäki , Henrik Boström , Panagiotis Papapetrou

We define the notion of a well-clusterable data set combining the point of view of the objective of $k$-means clustering algorithm (minimising the centric spread of data elements) and common sense (clusters shall be separated by gaps). We…

机器学习 · 计算机科学 2020-04-07 Mieczysław A. Kłopotek

A measure of distance between two clusterings has important applications, including clustering validation and ensemble clustering. Generally, such distance measure provides navigation through the space of possible clusterings. Mostly used…

社会与信息网络 · 计算机科学 2015-09-01 Reihaneh Rabbany , Osmar R. Zaïane

Typical Bayesian methods for models with latent variables (or random effects) involve directly sampling the latent variables along with the model parameters. In high-level software code for model definitions (using, e.g., BUGS, JAGS, Stan),…

统计计算 · 统计学 2022-12-12 E. C. Merkle , D. Furr , S. Rabe-Hesketh