Related papers: The probability distribution as a computational re…
Agglomerative hierarchical clustering is one of the most widely used approaches for exploring how observations in a dataset relate to each other. However, its greedy nature makes it highly sensitive to small perturbations in the data, often…
Large-scale multiple testing problems require the simultaneous assessment of many p-values. This paper compares several methods to assess the evidence in multiple binomial counts of p-values: the maximum of the binomial counts after…
We discuss several tests for whether a given set of independent and identically distributed (i.i.d.) draws does not come from a specified probability density function. The most commonly used are Kolmogorov-Smirnov tests, particularly…
Testing whether the observed data conforms to a purported model (probability distribution) is a basic and fundamental statistical task, and one that is by now well understood. However, the standard formulation, identity testing, fails to…
Let $b(x)$ be the probability that a sum of independent Bernoulli random variables with parameters $p_1, p_2, p_3, \ldots \in [0,1)$ equals $x$, where $\lambda := p_1 + p_2 + p_3 + \cdots$ is finite. We prove two inequalities for the…
A benefit of randomized experiments is that covariate distributions of treatment and control groups are balanced on average, resulting in simple unbiased estimators for treatment effects. However, it is possible that a particular…
We consider the computability of entropy and information in classical Hamiltonian systems. We define the information part and total information capacity part of entropy in classical Hamiltonian systems using relative information under a…
A collection of $n$ random events is said to be $(n - 1)$-wise independent if any $n - 1$ events among them are mutually independent. We characterise all probability measures with respect to which $n$ random events are $(n - 1)$-wise…
As a part of the construction of an information theory based on general probabilistic theories, we propose and investigate the several distinguishability measures and "entropies" in general probabilistic theories. As their applications,…
There is a random variable (X) with a determined outcome (i.e., X = x0), p(x0) = 1. Consider x0 to have a discrete uniform distribution over the integer interval [1, s], where the size of the sample space (s) = 1, in the initial state, such…
Computational complexity characterizes the usage of spatial and temporal resources by computational processes. In the classical theory of computation, e.g. in the Turing Machine model, computational processes employ only local space and…
We consider the structural change in a class of discrete valued time series that the conditional distribution follows a one-parameter exponential family. We propose a change-point test based on the maximum likelihood estimator of the…
Given samples from two non-negative random variables, we propose a family of tests for the null hypothesis that one random variable stochastically dominates the other at the second order. Test statistics are obtained as functionals of the…
We study generalizations of Demuth's Theorem, which states that the image of a Martin-L\"of random real under a tt-reduction is either computable or Turing equivalent to a Martin-L\"of random real. We show that Demuth's Theorem holds for…
Sequential estimation of a probability $p$ by means of inverse binomial sampling is considered. For $\mu_1,\mu_2>1$ given, the accuracy of an estimator $\hat{p}$ is measured by the confidence level $P[p/\mu_2\leq\hat{p}\leq p\mu_1]$. The…
Statistical machine learning theory often tries to give generalization guarantees of machine learning models. Those models naturally underlie some fluctuation, as they are based on a data sample. If we were unlucky, and gathered a sample…
This paper presents a procedure for testing the hypothesis that the underlying distribution of the data is elliptical when using robust location and scatter estimators instead of the sample mean and covariance matrix. Under mild assumptions…
A random set is a generalisation of a random variable, i.e. a set-valued random variable. The random set theory allows a unification of other uncertainty descriptions such as interval variable, mass belief function in Dempster-Shafer theory…
We want to select the best systems out of a given set of systems (or rank them) with respect to their expected performance. The systems allow random observations only and we assume that the joint observation of the systems has a…
Experiments often yield non-identically distributed data for statistical analysis. Tests of hypothesis under such set-ups are generally performed using the likelihood ratio test, which is non-robust with respect to outliers and model…