Related papers: Measuring the influence of the k-th largest variab…
The Sobol' indices are a recognized tool in global sensitivity analysis. When the uncertain variables in a model are statistically independent, the Sobol' indices may be easily interpreted and utilized. However, their interpretation and…
Linearization methods are customarily adopted in sampling surveys to obtain approximated variance formulae for estimators of nonlinear functions of finite population totals - such as ratios, correlation coefficients or measures of income…
Given a graph $G$, a community structure $\mathcal{C}$, and a budget $k$, the fair influence maximization problem aims to select a seed set $S$ ($|S|\leq k$) that maximizes the influence spread while narrowing the influence gap between…
The effect of the scalar spectral index on inflationary super-Hubble waves is to amplify/damp large wavelengths according to whether the spectrum is red ($n_{s}<1$) or blue ($n_{s}>1$). As a consequence, the large-scale temperature…
Clustering is a fundamental unsupervised learning task with applications across a wide range of domains. Popular algorithms such as $k$-means are efficient and widely used, but can be sensitive to outliers, ambiguous boundary points, and…
A coefficient is introduced that quantifies the extent of separation of a random variable $Y$ relative to a number of variables $\mathbf{X} = (X_1, \dots, X_p)$ by skillfully assessing the sensitivity of the relative effects of the…
This chapter covers different approaches to policy evaluation for assessing the causal effect of a treatment or intervention on an outcome of interest. As an introduction to causal inference, the discussion starts with the experimental…
For the first time, we introduce "Scaling invariable Benford distance" and "Benford cyclic graph", which can be used to analyze any data set. Using the quantity and the graph, we analyze some date sets with common distributions, such as…
We analyze fluctuations of random walks with generally distributed increments. Integral representations for key performance measures are obtained by extending an inversion theorem of Hewitt [11] for Laplace-Stieltjes transforms. Another…
We estimate the average of any arithmetic function $k$ over the values of any smooth polynomial in many variables provided only that $k$ has a distribution in arithmetic progressions of fixed modulus. We give several applications of this…
We revisit a double-scaled limit of the superconformal index of ${\cal N}=2$ superconformal field theories (SCFTs) which generalizes the Schur index. The resulting partition function, $\hat {\cal Z}(q,\alpha)$, has a standard $q$-expansion…
We provide an inferential framework to assess variable importance for heterogeneous treatment effects. This assessment is especially useful in high-risk domains such as medicine, where decision makers hesitate to rely on black-box treatment…
We introduce $k$-variance, a generalization of variance built on the machinery of random bipartite matchings. $K$-variance measures the expected cost of matching two sets of $k$ samples from a distribution to each other, capturing local…
A dataset has been classified by some unknown classifier into two types of points. What were the most important factors in determining the classification outcome? In this work, we employ an axiomatic approach in order to uniquely…
The hierarchically orthogonal functional decomposition of any measurable function f of a random vector X=(X_1,...,X_p) consists in decomposing f(X) into a sum of increasing dimension functions depending only on a subvector of X. Even when…
The h-index -- the value for which an individual has published at least h papers with at least h citations -- has become a popular metric to assess the citation impact of scientists. As already noted in the original work of Hirsch and as…
Influence diagnostics such as influence functions and approximate maximum influence perturbations are popular in machine learning and in AI domain applications. Influence diagnostics are powerful statistical tools to identify influential…
Based on the total integrability we first define an integral of a real valued function f as an interval function associated to its antiderivative F. By introducing the concept of the residue of a function into the real analysis, the…
We consider a random field $\phi(\mathbf{r})$ in $d$ dimensions which is largely concentrated around small `hotspots', with `weights', $w_i$. These weights may have a very broad distribution, such that their mean does not exist, or else is…
We define the $Q$-factor in the percolation problem as the quotient of the size of the largest cluster and the average size of all clusters. As the occupation probability $p$ is increased, the $Q$-factor for the system size $L$ grows…