English
Related papers

Related papers: Combining individually valid and conditionally i.i…

200 papers

A common task in high-throughput biology is to screen for associations across thousands of units of interest, e.g., genes or proteins. Often, the data for each unit are modeled as Gaussian measurements with unknown mean and variance and are…

Statistics Theory · Mathematics 2024-10-01 Nikolaos Ignatiadis , Bodhisattva Sen

The paper provides a simple test for deciding, from a given causal diagram, whether two sets of variables have the same bias-reducing potential under adjustment. The test requires that one of the following two conditions holds: either (1)…

Methodology · Statistics 2012-03-19 Judea Pearl , Azaria Paz

Estimation of the complete distribution of a random variable is a useful primitive for both manual and automated decision making. This problem has received extensive attention in the i.i.d. setting, but the arbitrary data dependent setting…

Machine Learning · Statistics 2023-03-01 Paul Mineiro , Steven R. Howard

We study least squares linear regression over $N$ uncorrelated Gaussian features that are selected in order of decreasing variance. When the number of selected features $p$ is at most the sample size $n$, the estimator under consideration…

Statistics Theory · Mathematics 2019-10-04 Ji Xu , Daniel Hsu

We prove that all polynomials in several variables can be decomposed as the sums of $k$th powers: $P(x_1,...,x_n) = Q_1(x_1,...,x_n)^k+...+ Q_s(x_1,...,x_n)^k$, provided that elements of the base field are themselves sums of $k$th powers.…

Number Theory · Mathematics 2011-10-20 Arnaud Bodin , Mireille Car

In many scientific contexts, different investigators experiment with or observe different variables with data from a domain in which the distinct variable sets might well be related. This sort of fragmentation sometimes occurs in molecular…

Artificial Intelligence · Computer Science 2019-09-05 Shuyan Wang

In practice functional data are sampled on a discrete set of observation points and often susceptible to noise. We consider in this paper the setting where such data are used as explanatory variables in a regression problem. If the primary…

Methodology · Statistics 2021-12-14 Siegfried Hörmann , Fatima Jammoul

The problem of detecting changes in covariance for a single pair of features has been studied in some detail, but may be limited in importance or general applicability. In contrast, testing equality of covariance matrices of a {\it set} of…

Methodology · Statistics 2017-12-12 Yi-Hui Zhou

There is currently a great deal of interest in the 4D-Var data assimilation scheme, in which one uses observational data to find the optimal initial condition for a differential equation by minimizing a cost function over the set of all…

Optimization and Control · Mathematics 2012-07-20 Graham Cox

A question that comes up repeatedly is how to combine the results of two experiments if all that is known is that one experiment had a n-sigma effect and another experiment had a m-sigma effect. This question is not well-posed: depending on…

Data Analysis, Statistics and Probability · Physics 2008-12-20 Robert D. Cousins

In Bayesian hypothesis testing, evidence for a statistical model is quantified by the Bayes factor, which represents the relative likelihood of observed data under that model compared to another competing model. In general, computing Bayes…

Computation · Statistics 2021-12-07 Thomas J. Faulkenberry

In an attempt to provide an answer to the increasing criticism against p-values and to bridge the gap between statistical inference and prediction modelling, we introduce the probability of improved prediction (PIP). In general, the PIP is…

Methodology · Statistics 2024-05-28 Olivier Thas , Stijn Jaspers

Prediction, where observed data is used to quantify uncertainty about a future observation, is a fundamental problem in statistics. Prediction sets with coverage probability guarantees are a common solution, but these do not provide…

Statistics Theory · Mathematics 2022-11-22 Leonardo Cella , Ryan Martin

For testing two random vectors for independence, we consider testing whether the distance of one vector from a center point is independent from the distance of the other vector from a center point by a univariate test. In this paper we…

Methodology · Statistics 2016-03-11 Ruth Heller , Yair Heller

Positive predictive value and negative predictive value are two widely used parameters to assess the clinical usefulness of a medical diagnostic test. When there are two diagnostic tests, it is recommendable to make a comparative assessment…

Methodology · Statistics 2024-05-29 Antonio Martín Andrés , Pedro Femia Marzo

Assessing the statistical significance of parameter estimates is an important step in high-dimensional vector autoregression modeling. Using the least-squares boosting method, we compute the p-value for each selected parameter at every…

Econometrics · Economics 2023-03-16 Xiao Huang

Identifying effects of actions (treatments) on outcome variables from observational data and causal assumptions is a fundamental problem in causal inference. This identification is made difficult by the presence of confounders which can be…

Methodology · Statistics 2012-03-19 Ilya Shpitser , Tyler VanderWeele , James M. Robins

This paper concerns the probabilistic evaluation of the effects of actions in the presence of unmeasured variables. We show that the identification of causal effect between a singleton variable X and a set of variables Y can be accomplished…

Artificial Intelligence · Computer Science 2013-02-21 David Galles , Judea Pearl

The problem of recovering (count and sum) range queries over multidimensional data only on the basis of aggregate information on such data is addressed. This problem can be formalized as follows. Suppose that a transformation T producing a…

Databases · Computer Science 2007-05-23 Francesco Buccafurri , Filippo Furfaro , Domenico Sacca'

A classifier for two or more samples is proposed when the data are high-dimensional and the underlying distributions may be non-normal. The classifier is constructed as a linear combination of two easily computable and interpretable…

Statistics Theory · Mathematics 2016-08-02 M. Rauf Ahmad , Tatjana Pavlenko