Related papers: Records from partial comparisons and discrete appr…
This paper focuses on nonparametric statistical inference of the hazard rate function of discrete distributions based on $\delta$-record data. We derive the explicit expression of the maximum likelihood estimator and determine its exact…
In this paper, we present methods of obtaining single moments of order statistics arising from posibly dependent and non-identically distributed discrete random variables. We derive exact and approximate formulas convenient for numerical…
Testing for conditional independence is a core aspect of constraint-based causal discovery. Although commonly used tests are perfect in theory, they often fail to reject independence in practice, especially when conditioning on multiple…
Marginal imputation, which consists of imputing each item requiring imputation separately, is often used in surveys. This type of imputation procedures leads to asymptotically unbiased estimators of simple parameters such as population…
This paper investigates what can be inferred about an arbitrary continuous probability distribution from a finite sample of $N$ observations drawn from it. The central finding is that the $N$ sorted sample points partition the real line…
We investigate the problems of identity and closeness testing over a discrete population from random samples. Our goal is to develop efficient testers while guaranteeing Differential Privacy to the individuals of the population. We describe…
Let $\mathbf{X}(n) \in \mathbb{R}^d$ be a sequence of random vectors, where $n\in\mathbb{N}$ and $d = d(n)$. Under certain weakly dependence conditions, we prove that the distribution of the maximal component of $\mathbf{X}$ and the…
Regular variation provides a convenient theoretical framework to study large events. In the multivariate setting, the dependence structure of the positive extremes is characterized by a measure - the spectral measure - defined on the…
Two closely related discrete probability distributions are introduced. In each case the support is a set of vectors in $\mathbb{R}^n$ obtained from the partitions of the fixed positive integer $n$. These distributions arise naturally when…
Statistical matching is a technique for integrating two or more data sets when information available for matching records for individual participants across data sets is incomplete. Statistical matching can be viewed as a missing data…
In the context of this paper, a record is an entry in a sequence of random variables (RV's) that is larger or smaller than all previous entries. After a brief review of the classic theory of records, which is largely restricted to sequences…
While supporting the execution of business processes, information systems record event logs. Conformance checking relies on these logs to analyze whether the recorded behavior of a process conforms to the behavior of a normative…
The estimation of information measures of continuous distributions based on samples is a fundamental problem in statistics and machine learning. In this paper, we analyze estimates of differential entropy in $K$-dimensional Euclidean space,…
We consider two variables that are related to each other by an invertible function. While it has previously been shown that the dependence structure of the noise can provide hints to determine which of the two variables is the cause, we…
We establish properties of a new type of fractal which has partial self similarity at all scales. For any collection of iterated functions systems with an associated probability distribution and any positive integer V there is a…
The identification of increasingly smaller signal from objects observed with a non-perfect instrument in a noisy environment poses a challenge for a statistically clean data analysis. We want to compute the probability of frequencies…
Partial correlations quantify linear association between two variables adjusting for the influence of the remaining variables. They form the backbone for graphical models and are readily obtained from the inverse of the covariance matrix.…
We review recent advances on the record statistics of strongly correlated time series, whose entries denote the positions of a random walk or a L\'evy flight on a line. After a brief survey of the theory of records for independent and…
This work considers a problem of estimating a mixing probability density $f$ in the setting of discrete mixture models. The paper consists of three parts. The first part focuses on the construction of an $L_1$ consistent estimator of $f$.…
Distance correlation is a measure of dependence between two paired random vectors or matrices of arbitrary, not necessarily equal, dimensions. Unlike Pearson correlation, the population distance correlation coefficient is zero if and only…