Related papers: Knockoffs for exchangeable categorical covariates
The knockoffs is a recently proposed powerful framework that effectively controls the false discovery rate (FDR) for variable selection. However, none of the existing knockoff solutions are directly suited to handle multivariate or…
The aim of this paper is to establish Hoeffding and Bernstein type concentration inequalities for weighted sums of exchangeable random variables. A special case is the i.i.d. setting, where random variables are sampled independently from…
A sequence of random variables is called exchangeable if the joint distribution of the sequence is unchanged by any permutation of the indices. De Finetti's theorem characterizes all $\{0,1\}$-valued exchangeable sequences as a "mixture" of…
Model-X knockoffs is a flexible wrapper method for high-dimensional regression algorithms, which provides guaranteed control of the false discovery rate (FDR). Due to the randomness inherent to the method, different runs of model-X…
Let $(X_1,\ldots,X_n)$ be an exchangeable random vector with distribution function $F$, and denote by $Y_1\leq \cdots\leq Y_n$ the corresponding order statistics. We show that the conditional distribution of $(X_1,\ldots,X_n)$ given…
Covariate shift, a widely used assumption in tackling {\it distributional shift} (when training and test distributions differ), focuses on scenarios where the distribution of the labels conditioned on the feature vector is the same, but the…
We explore the class of exchangeable Bernoulli distributions building on their geometrical structure. Exchangeable Bernoulli probability mass functions are points in a convex polytope and we have found analytical expressions for their…
A new statistical procedure (Model-X \cite{candes2018}) has provided a way to identify important factors using any supervised learning method controlling for FDR. This line of research has shown great potential to expand the horizon of…
We establish Hoeffding-type concentration inequalities for the low and high tail bounds of sums of exchangeable random variables. Our results exhibit an anti-symmetry in such tail bounds due to the assumption of exchangeability, a…
A $\widetilde{Q}-$representation of real numbers is introduced as a generalization of the $p-$adic and $Q-$representations. It is shown that the $\widetilde{Q}-$representation may be used as a convenient tool for the construction and study…
Model-free knockoffs is a recently proposed technique for identifying covariates that is likely to have an effect on a response variable. The method is an efficient method to control the false discovery rate in hypothesis tests for separate…
The goal of feature selection is to identify important features that are relevant to explain an outcome variable. Most of the work in this domain has focused on identifying globally relevant features, which are features that are related to…
We introduce a general framework for de Finetti reduction results, applicable to various notions of partially exchangeable probability distributions. Explicit statements are derived for the cases of exchangeability, Markov exchangeability,…
Recently, the scheme of model-X knockoffs was proposed as a promising solution to address controlled feature selection under high-dimensional finite-sample settings. However, the procedure of model-X knockoffs depends heavily on the…
This paper develops a method based on model-X knockoffs to find conditional associations that are consistent across diverse environments, controlling the false discovery rate. The motivation for this problem is that large data sets may…
The hitting partitions are random partitions that arise from the investigation of so-called hitting scenarios of max-infinitely-divisible (max-i.d.)~distributions. We study a class of max-i.d.~laws with exchangeable hitting partitions…
We study the independence structure of finitely exchangeable distributions over random vectors and random networks. In particular, we provide necessary and sufficient conditions for an exchangeable vector so that its elements are completely…
We study finite-sample inference for the trade-off function of two unknown probability distributions, the function that traces the optimal type I/type II error frontier in binary testing. Given samples from distributions $P$ and $Q$, we…
Interpretability and stability are two important features that are desired in many contemporary big data applications arising in economics and finance. While the former is enjoyed to some extent by many existing forecasting approaches, the…
Controlling false discovery rate (FDR) is crucial for variable selection, multiple testing, among other signal detection problems. In literature, there is certainly no shortage of FDR control strategies when selecting individual features,…