Related papers: Records from partial comparisons and discrete appr…
We introduce the coverage correlation coefficient, a novel nonparametric measure of statistical association designed to quantifies the extent to which two random variables have a joint distribution concentrated on a singular subset with…
Let $X_1,~X_2,\cdots$ be a sequence of i.i.d random variables which are supposed to be observed in sequence. The $n$th value in the sequence is a $k-record~value$ if exactly $k$ of the first $n$ values (including $X_n$) are at least as…
We study the statistics of the number of records $R_n$ for a symmetric, $n$-step, discrete jump process on a $1D$ lattice. At a given step, the walker can jump by arbitrary lattice units drawn from a given symmetric probability…
We propose a new estimation procedure of the conditional density for independent and identically distributed data. Our procedure aims at using the data to select a function among arbitrary (at most countable) collections of candidates. By…
In this paper we examine some relative orderings of upper and lower records. It is shown that if m > n, the mth upper record ages faster than the nth upper record, where the data sets come from a sequence of independent and identically…
In this paper, we systematically summarize and enhance the understanding of weak convergence and functional limits of record numbers in discrete-time random walks under Spitzer's condition, and extend these findings to $\sigma$--record…
We obtain the posterior distribution of a random process conditioned on observing the empirical frequencies of a finite sample path. We find under a rather broad assumption on the "dependence structure" of the process, {\em c.f.}…
This article develops a method to construct the optimal sequential test for monitoring the changes in the distribution of finite observation sequences with a general dependence structure. This method allows us to prove that different…
In machine learning, the performance of a classifier depends on both the classifier model and the separability/complexity of datasets. To quantitatively measure the separability of datasets, we create an intrinsic measure -- the…
This version is ***superseded*** by a full version that can be found at http://www.itu.dk/people/pagh/papers/mining-jour.pdf, which contains stronger theoretical results and fixes a mistake in the reporting of experiments. Abstract:…
We study an online version of the noisy binary search problem where feedback is generated by a non-stochastic adversary rather than perturbed by random noise. We reframe this as maintaining an accurate estimate for the median of an…
We consider a type of pull voting suitable for discrete numeric opinions which can be compared on a linear scale, for example, 1 ('disagree strongly'), 2 ('disagree'), $\ldots,$ 5 ('agree strongly'). On observing the opinion of a random…
We study properties of two resampling scenarios: Conditional Randomisation and Conditional Permutation schemes, which are relevant for testing conditional independence of discrete random variables $X$ and $Y$ given a random variable $Z$.…
Comparing the differences in outcomes (that is, in "dependent variables") between two subpopulations is often most informative when comparing outcomes only for individuals from the subpopulations who are similar according to "independent…
Density-based directed distances -- particularly known as divergences -- between probability distributions are widely used in statistics as well as in the adjacent research fields of information theory, artificial intelligence and machine…
It is shown that statistics of records for time series generated by random walks are independent of the details of the jump distribution, as long as the latter is continuous and symmetric. In N steps, the mean of the record distribution…
We derive conditions under which random sequences of polarizations (two-point symmetrizations) converge almost surely to the symmetric decreasing rearrangement. The parameters for the polarizations are independent random variables whose…
It is common to conduct causal inference in matched observational studies by proceeding as though treatment assignments within matched sets are assigned uniformly at random and using this distribution as the basis for inference. This…
The study of "random segments" is a classic issue in geometrical probability, whose complexity depends on how it is defined. But in apparently simple models, the random behavior is not immediate. In the present manuscript the following…
We are interested in learning causal relationships between pairs of random variables, purely from observational data. To effectively address this task, the state-of-the-art relies on strong assumptions regarding the mechanisms mapping…