Related papers: Knockoffs for exchangeable categorical covariates
Continuous improvement in medical imaging techniques allows the acquisition of higher-resolution images. When these are used in a predictive setting, a greater number of explanatory variables are potentially related to the dependent…
The Model-X knockoffs is a practical methodology for variable selection, which stands out from other selection strategies since it allows for the control of the false discovery rate (FDR), relying on finite-sample guarantees. In this…
Let $X_{d_1,d_2}$ be an $F$-random variable with numerator and denominator degrees of freedom $d_1$ and $d_2$, respectively. We investigate the inequality: $P\{|X_{d_1,d_2}-E[X_{d_1,d_2}]|\le \sqrt{{\rm Var}(X_{d_1,d_2})}\}\ge…
We introduce an iterative discrete information production process where we can extend ordered normalised vectors by new elements based on a simple affine transformation, while preserving the predefined level of inequality, G, as measured by…
We consider the problem of variable selection in regression models. In particular, we are interested in selecting explanatory covariates linked with the response variable and we want to determine which covariates are relevant, that is which…
A multivariate distribution function F is in the max-domain of attraction of an extreme value distribution if and only if this is true for the copula corresponding to F and its univariate margins. Aulbach et al. (2012a) have shown that a…
We study an information analogue of infinitely divisible probability distributions, where the i.i.d. sum is replaced by the joint distribution of an i.i.d. sequence. A random variable $X$ is called informationally infinitely divisible if,…
Let (X_n) be a sequence of random variables (with values in a separable metric space) and (N_n) a sequence of random indices. Conditions for X_{N_n} to converge stably (in particular, in distribution) are provided. Some examples, where such…
The roundoff errors in computer simulations of continuous dynamical systems, caused by finiteness of machine arithmetic, can lead to qualitative discrepancies between phase portraits of the resulting spatially discretized systems and the…
We consider problems where many, somewhat redundant, hypotheses are tested and we are interested in reporting the most precise rejections, with false discovery rate (FDR) control. This is the case, for example, when researchers are…
Consider a set of multivariate distributions, $F_1,\dots,F_M$, aiming to explain the same phenomenon. For instance, each $F_m$ may correspond to a different candidate background model for calibration data, or to one of many possible signal…
In this paper, we study the asymptotic behavior of the extreme eigenvalues and eigenvectors of the high dimensional spiked sample covariance matrices, in the supercritical case when a reliable detection of spikes is possible. Especially, we…
A collaborative distributed binary decision problem is considered. Two statisticians are required to declare the correct probability measure of two jointly distributed memoryless process, denoted by $X^n=(X_1,\dots,X_n)$ and…
Partially exchangeable sequences representable as mixtures of Markov chains are completely specified by de Finetti's mixing measure. The paper characterizes, in terms of a subclass of hidden Markov models, the partially exchangeable…
Given a sequence \xi_1, \xi_2,... of X-valued, exchangeable random elements, let q(\xi^(n)) and p_m(\xi^(n)) stand for posterior and predictive distribution, respectively, given \xi^(n) = (\xi_1,..., \xi_n). We provide an upper bound for…
We provide new non-asymptotic false discovery proportion (FDP) confidence envelopes in several multiple testing settings relevant for modern high dimensional-data methods. We revisit the multiple testing scenarios considered in the recent…
It is well known that the entropy $H(X)$ of a discrete random variable $X$ is always greater than or equal to the entropy $H(f(X))$ of a function $f$ of $X$, with equality if and only if $f$ is one-to-one. In this paper, we give tight…
Conditional independence testing (CIT) is essential for reliable scientific discovery. It prevents spurious findings and enables controlled feature selection. Recent CIT methods have used machine learning (ML) models as surrogates of the…
In modern scientific research, the objective is often to identify which variables are associated with an outcome among a large class of potential predictors. This goal can be achieved by selecting variables in a manner that controls the the…
We consider the problem of inference in shift-share research designs. The choice between existing approaches that allow for unrestricted spatial correlation involves tradeoffs, varying in terms of their validity when there are relatively…