Related papers: E-Statistics, Group Invariance and Anytime Valid T…
A rank-invariant clustering of variables is introduced that is based on the predictive strength between groups of variables, i.e., two groups are assigned a high similarity if the variables in the first group contain high predictive…
We study the inference problem in the group testing to identify defective items from the perspective of the decision theory. We introduce Bayesian inference and consider the Bayesian optimal setting in which the true generative process of…
In the Bayesian literature on model comparison, Bayes factors play the leading role. In the classical statistical literature, model selection criteria are often devised used cross-validation ideas. Amalgamating the ideas of Bayes factor and…
A significant obstacle in the development of robust machine learning models is covariate shift, a form of distribution shift that occurs when the input distributions of the training and test sets differ while the conditional label…
Undirected graphical models encode in a graph $G$ the dependency structure of a random vector $Y$. In many applications, it is of interest to model $Y$ given another random vector $X$ as input. We refer to the problem of estimating the…
We propose a novel infection spread model based on a random connection graph which represents connections between $n$ individuals. Infection spreads via connections between individuals and this results in a probabilistic cluster formation…
Gaussian graphical models with sparsity in the inverse covariance matrix are of significant interest in many modern applications. For the problem of recovering the graphical structure, information criteria provide useful optimization…
A central goal of machine learning is to learn robust representations that capture the causal relationship between inputs features and output labels. However, minimizing empirical risk over finite or biased datasets often results in models…
In the recent Basel Accords, the Expected Shortfall (ES) replaces the Value-at-Risk (VaR) as the standard risk measure for market risk in the banking sector, making it the most important risk measure in financial regulation. One of the most…
The problem of selecting optimal backdoor adjustment sets to estimate causal effects in graphical models with hidden and conditioned variables is addressed. Previous work has defined optimality as achieving the smallest asymptotic…
We study extreme values of group-indexed stable random fields for discrete groups $G$ acting geometrically on spaces $X$ in the following cases: 1) $G$ acts freely, properly discontinuously by isometries on a CAT(-1) space $X$, 2) $G$ is a…
Let $(G,\rho)$ be a stationary random graph, and use $B^G_{\rho}(r)$ to denote the ball of radius $r$ about $\rho$ in $G$. Suppose that $(G,\rho)$ has annealed polynomial growth, in the sense that $\mathbb{E}[|B^G_{\rho}(r)|] \leq O(r^k)$…
We analyze common types of e-variables and e-processes for composite exponential family nulls: the optimal e-variable based on the reverse information projection (RIPr), the conditional (COND) e-variable, and the universal inference (UI)…
The enhanced power graph $\mathcal{P}_e(G)$ of a group $G$ is a graph with vertex set $G$ and two vertices are adjacent if they belong to the same cyclic subgroup. In this paper, we consider the minimum degree, independence number and…
The empirical risk minimization approach to data-driven decision making requires access to training data drawn under the same conditions as those that will be faced when the decision rule is deployed. However, in a number of settings, we…
In this paper we study the action of a countable group $\Gamma$ on the space of orders on the group. In particular, we are concerned with the invariant probability measures on this space, known as invariant random orders. We show that for…
A non parametric method based on the empirical likelihood is proposed for detecting the change in the coefficients of high-dimensional linear model where the number of model variables may increase as the sample size increases. This amounts…
In big data analysis for detecting rare and weak signals among $n$ features, some grouping-test methods such as Higher Criticism test (HC), Berk-Jones test (B-J), and $\phi$-divergence test share the similar asymptotical optimality when $n…
Consider $n$ iid real-valued random vectors of size $k$ having iid coordinates with a general distribution function $F$. A vector is a maximum if and only if there is no other vector in the sample which weakly dominates it in all…
We study the conditional distribution of goodness of fit statistics of the Cram\'{e}r--von Mises type given the complete sufficient statistics in testing for exponential family models. We show that this distribution is close, in large…