Related papers: Measuring the influence of the k-th largest variab…
We introduce an index for measuring the influence of the k-th smallest variable on a pseudo-Boolean function. This index is defined from a weighted least squares approximation of the function by linear combinations of order statistic…
Consider a Boolean function f on the n-dimensional hypercube, and a set of variables (indexed by) $S \subset \{1,2,\ldots,n\}.$ The coalition influence of the variables S on a function f is the probability that after a random assignment of…
Sobol' indices measure the dependence of a high dimensional function on groups of variables defined on the unit cube $[0,1]^d$. They are based on the ANOVA decomposition of functions, which is an $L^2$ decomposition. In this paper we…
We study the problem of estimating a monotone function $f:\{0,1\}^d\to[0,1]$ from noisy observations at uniformly random vertices of the Boolean hypercube. As a measure of complexity for the target~$f$, we use the total $L^1$-influence…
For a function $f$ over the discrete cube, the total $L_1$ influence of $f$ is defined as $\sum_{i=1}^n \|\partial_i f\|_1$, where $\partial_i f$ denotes the discrete derivative of $f$ in the direction $i$. In this work, we show that the…
The K function and its related statistics have been an enduring tool in the analysis of spatial point processes, providing an easy to compute and interpret summary statistic for characterising the interactions between points of one type, or…
Variable selection, also known as feature selection in machine learning, plays an important role in modeling high dimensional data and is key to data-driven scientific discoveries. We consider here the problem of detecting influential…
Inspired by the $k$-inversion statistic for LLT polynomials, we define a $k$-inversion number and $k$-descent set for words. Using these, we define a new statistic on words, called the $k$-major index, that interpolates between the major…
Composite indicators aggregate a set of variables using weights which are understood to reflect the variables' importance in the index. In this paper we propose to measure the importance of a given variable within existing composite…
In this paper we consider the influences of variables on Boolean functions in general product spaces. Unlike the case of functions on the discrete cube where there is a clear definition of influence, in the general case at least three…
The covariance between real finite variance random variables can be expressed as the commutator of taking expectations and multiplying, both viewed as operators extended to act jointly on pairs of functions. The efficient influence curve of…
Effect size indices are useful tools in study design and reporting because they are unitless measures of association strength that do not depend on sample size. Existing effect size indices are developed for particular parametric models or…
Influence diagnosis is important since presence of influential observations could lead to distorted analysis and misleading interpretations. For high-dimensional data, it is particularly so, as the increased dimensionality and complexity…
A variety of bibliometric measures have been proposed to quantify the impact of researchers and their work. The h-index is a notable and widely-used example which aims to improve over simple metrics such as raw counts of papers or…
When trying to gain better visibility into a machine learning model in order to understand and mitigate the associated risks, a potentially valuable source of evidence is: which training examples most contribute to a given behavior?…
Estimators based on influence functions (IFs) have been shown to be effective in many settings, especially when combined with machine learning techniques. By focusing on estimating a specific target of interest (e.g., the average effect of…
An index transform, involving the square of Whittaker's function is introduced and investigated. The corresponding inversion formula is established. Particular cases cover index transforms of the Lebedev type with products of the modified…
We use a relative trace formula on GL(2) to compute a sum of twisted modular L-functions anywhere in the critical strip, weighted by a Fourier coefficient and a Hecke eigenvalue. When the weight k or level N is sufficiently large, the sum…
We propose a coefficient that measures the dependence among large values for spatial processes of maxima. Its main properties are: a) $k$ locations can be taken into account; b) it takes values in $[0,1]$ and higher values indicate stronger…
A class of Fourier based statistics for irregular spaced spatial data is introduced, examples include, the Whittle likelihood, a parametric estimator of the covariance function based on the $L_{2}$-contrast function and a simple…