Related papers: A Note On k-Means Probabilistic Poverty
We provide a permutation-invariant version of the Koml\'os' theorem for non-negative random variables. The proof is quite elementary in the sense that it did not use the Axiom of Choice, and was based on a recent result in [3].
We provide necessary and sufficient conditions for the uniqueness of the k-means set of a probability distribution. This uniqueness problem is related to the choice of k: depending on the underlying distribution, some values of this…
Following Fisher, it is widely believed that randomization "relieves the experimenter from the anxiety of considering innumerable causes by which the data may be disturbed." In particular, it is said to control for known and unknown…
We firstly show that the standard interpretation of natural quantification in mathematical logic does not provide a satisfying account of its original richness. In particular, it ignores the difference between generic and distributive…
The widely applied k-means algorithm produces clusterings that violate our expectations with respect to high/low similarity/density and is in conflict with Kleinberg's axiomatic system for distance based clustering algorithms that…
The Kochen-Specker theorem states that exclusive and complete deterministic outcome assignments are impossible for certain sets of measurements, called Kochen-Specker (KS) sets. A straightforward consequence is that KS sets do not have…
In random expected utility (Gul and Pesendorfer, 2006), the distribution of preferences is uniquely recoverable from random choice. This paper shows through two examples that such uniqueness fails in general if risk preferences are random…
Random partition models are widely used in Bayesian methods for various clustering tasks, such as mixture models, topic models, and community detection problems. While the number of clusters induced by random partition models has been…
Autoregressive language models (LMs) map token sequences to probabilities. The usual practice for computing the probability of any character string (e.g. English sentences) is to first transform it into a sequence of tokens that is scored…
We show that the phenomenon of anomalous weak values is not limited to quantum theory. In particular, we show that the same features occur in a simple model of a coin subject to a form of classical backaction with pre- and post-selection.…
Probabilistic metrology attempts to improve parameter estimation by occasionally reporting an excellent estimate and the rest of the time either guessing or doing nothing at all. Here we show that probabilistic metrology can never improve…
The $k$-means algorithm is a prevalent clustering method due to its simplicity, effectiveness, and speed. However, its main disadvantage is its high sensitivity to the initial positions of the cluster centers. The global $k$-means is a…
We derive the most probable distribution of resources for a simple society. We find that a probabilistic analysis forbids both too much and too less equity, and selects instead a minimally ordered state. We give the detailed calculations…
We show that irreducibility is not a first-order definable property of real algebraic varieties. The proof is based on the recent o-minimality result for the exponential function. We conjecture that irreducibility is not a definable…
We present an alternative approach to the Bayesian nonparametric analysis of conditional species richness under two-parameter Poisson Dirichlet priors. We rely on a known characterization by deletion of classes property and on results for…
We consider first order expressible properties of random perfect graphs. That is, we pick a graph $G_n$ uniformly at random from all (labelled) perfect graphs on $n$ vertices and consider the probability that it satisfies some graph…
For nonnegative random variables with finite means we introduce an analogous of the equilibrium residual-lifetime distribution based on the quantile function. This allows to construct new distributions with support (0,1), and to obtain a…
In this note, we show that classical statistical tests for randomness are language dependent.
We explain the concept of p-values presupposing only rudimentary probability theory. We also use the occasion to introduce the notion of p-function, so that p-values are values of a p-function. The explanation is restricted to the discrete…
There has been great interest in fairness in machine learning, especially in relation to classification problems. In ranking-related problems, such as in online advertising, recommender systems, and HR automation, much work on fairness…