Related papers: Cumulative differences between paired samples
One of the major challenges for collective intelligence is inconsistency, which is unavoidable whenever subjective assessments are involved. Pairwise comparisons allow one to represent such subjective assessments and to process them by…
This paper considers estimating a covariance matrix of $p$ variables from $n$ observations by either banding or tapering the sample covariance matrix, or estimating a banded version of the inverse of the covariance. We show that these…
Scientists frequently generalize population level causal quantities such as average treatment effect from a source population to a target population. When the causal effects are heterogeneous, differences in subject characteristics between…
Estimating the causal effect of a treatment or health policy with observational data can be challenging due to an imbalance of and a lack of overlap between treated and control covariate distributions. In the presence of limited overlap,…
It is well known that parameters for strongly correlated predictor variables in a linear model cannot be accurately estimated. We look for linear combinations of these parameters that can be. Under a uniform model, we find such linear…
In various disordered systems or non-equilibrium dynamical models, the large deviations of some observables have been found to display different scalings for rare values bigger or smaller than the typical value. In the present paper, we…
Weighting methods in causal inference have been widely used to achieve a desirable level of covariate balancing. However, the existing weighting methods have desirable theoretical properties only when a certain model, either the propensity…
The Curie-Weiss model, originally used to study phase transitions in statistical mechanics, has been adapted to model phenomena in social sciences where many agents interact with each other. Reconstructing the probability measure of a…
A statistical analysis of optimal universal cloning shows that it is possible to identify an ideal (but non-positive) copying process that faithfully maps all properties of the original Hilbert space onto two separate quantum systems. The…
In randomized experiments, treatment and control groups should be roughly the same--balanced--in their distributions of pretreatment variables. But how nearly so? Can descriptive comparisons meaningfully be paired with significance tests?…
The average distance from a node to all other nodes in a graph, or from a query point in a metric space to a set of points, is a fundamental quantity in data analysis. The inverse of the average distance, known as the (classic) closeness…
We propose a probabilistic model to aggregate the answers of respondents answering multiple-choice questions. The model does not assume that everyone has access to the same information, and so does not assume that the consensus answer is…
The network scale-up method enables researchers to estimate the size of hidden populations, such as drug injectors and sex workers, using sampled social network data. The basic scale-up estimator offers advantages over other size estimation…
Binary particle coagulation can be modelled as the repeated random process of the combination of two particles to form a third. The kinetics can be represented by population rate equations based on a mean field assumption, according to…
In time series analysis, statistics based on collections of estimators computed from sub-samples play a crucial role in an increasing variety of important applications. Proving results about the joint asymptotic distribution of such…
When estimating causal effects using observational data, it is desirable to replicate a randomized experiment as closely as possible by obtaining treated and control groups with similar covariate distributions. This goal can often be…
Terms such as leader, mediator, and follower sound equal in the description of a pack of wolves, a street protest crowd, or a business team and have very similar meanings. This indicates the presence of some general law or structure that…
In this manuscript a unified framework for conducting inference on complex aggregated data in high dimensional settings is proposed. The data are assumed to be a collection of multiple non-Gaussian realizations with underlying undirected…
The Wiener index (i.e., the total distance or the transmission number), defined as the sum of distances between all unordered pairs of vertices in a graph, is one of the most popular molecular descriptors. In this article we summarize some…
Silhouette coefficient is an established internal clustering evaluation measure that produces a score per data point, assessing the quality of its clustering assignment. To assess the quality of the clustering of the whole dataset, the…