Related papers: Measuring the influence of the k-th largest variab…
Let {(Z_i,W_i):i=1,...,n} be uniformly distributed in [0,1]^d * G(k,d), where G(k,d) denotes the space of k-dimensional linear subspaces of R^d. For a differentiable function f from [0,1]^k to [0,1]^d we say that f interpolates (z,w) in…
We consider the problem of learning a $d$-variate function $f$ defined on the cube $[-1,1]^d\subset {\mathbb R}^d$, where the algorithm is assumed to have black box access to samples of $f$ within this domain. Denote ${\mathcal S}_r \subset…
Machine learning systems such as large scale recommendation systems or natural language processing systems are usually trained on billions of training points and are associated with hundreds of billions or trillions of parameters. Improving…
The size of the effect of the difference in two groups with respect to a variable of interest may be estimated by the classical Cohen's $d$. A recently proposed generalized estimator allows conditioning on further independent variables…
Instrumental variable models allow us to identify a causal function between covariates $X$ and a response $Y$, even in the presence of unobserved confounding. Most of the existing estimators assume that the error term in the response $Y$…
Many existing approaches to generalizing statistical inference amidst distribution shift operate under the covariate shift assumption, which posits that the conditional distribution of unobserved variables given observable ones is invariant…
The factor analysis model is a statistical model where a certain number of hidden random variables, called factors, affect linearly the behaviour of another set of observed random variables, with additional random noise. The main assumption…
How does the training data affect a model's behavior? This is the question we seek to answer with data attribution. The leading practical approaches to data attribution are based on influence functions (IF). IFs utilize a first-order Taylor…
The multi-index model is a simple yet powerful high-dimensional regression model which circumvents the curse of dimensionality assuming $ \mathbb{E} [ Y | X ] = g(A^\top X) $ for some unknown index space $A$ and link function $g$. In this…
Under a general structural equation framework for causal inference, we provide a definition of the causal effect of a variable X on another variable Y, and propose an approach to estimate this causal effect via the use of instrumental…
The Hirsch function of a given continuous function is a new function depending on a parameter. It exists provided some assumptions are satisfied. If this parameter takes the value one, we obtain the well-known h-index. We prove some…
Inference is the task of drawing conclusions about unobserved variables given observations of related variables. Applications range from identifying diseases from symptoms to classifying economic regimes from price movements. Unfortunately,…
We consider the issue of assessing influence of observations in the class of Birnbaum-Saunders nonlinear regression models, which is useful in lifetime data analysis. Our results generalize those in Galea et al. [2004, Influence diagnostics…
Subsampling methods have been recently proposed to speed up least squares estimation in large scale settings. However, these algorithms are typically not robust to outliers or corruptions in the observed covariates. The concept of influence…
The {\em Total Influence} ({\em Average Sensitivity) of a discrete function is one of its fundamental measures. We study the problem of approximating the total influence of a monotone Boolean function \ifnum\plusminus=1 $f: \{\pm1\}^n…
We study, in a global uniform manner, the quotient of the ring of polynomials in l sets of n variables, by the ideal generated by diagonal quasi-invariant polynomials for general permutation groups W=G(r,n). We show that, for each such…
We conduct a KL-divergence based procedure for testing elliptical distributions. The procedure simultaneously takes into account the two defining properties of an elliptically distributed random vector: independence between length and…
Originally developed for measuring the heterogeneity of wealth measures, inequality indices are quantitative scores that take values in the unit interval, with the zero score characterizing perfect equality. In this paper, we draw attention…
Influence functions (IFs) elucidate how training data changes model behavior. However, the increasing size and non-convexity in large-scale models make IFs inaccurate. We suspect that the fragility comes from the first-order approximation…
Rate coefficients can fluctuate in statically and dynamically disordered kinetics. Here we relate the rate coefficient for an irreversibly decaying population to the Fisher information. From this relationship we define kinetic versions of…