English
Related papers

Related papers: Cumulative differences between paired samples

200 papers

Matching is a widely used causal inference design that aims to approximate a randomized experiment using observational data by forming matched sets of treated and control units based on similarities in their covariates. Ideally, treated…

Methodology · Statistics 2026-04-06 Jianan Zhu , Jeffrey Zhang , Zijian Guo , Siyu Heng

It has been proposed that complex populations, such as those that arise in genomics studies, may exhibit dependencies among observations as well as among variables. This gives rise to the challenging problem of analyzing unreplicated…

Machine Learning · Statistics 2018-06-08 Michael Hornstein , Roger Fan , Kerby Shedden , Shuheng Zhou

The quotient correlation is defined here as an alternative to Pearson's correlation that is more intuitive and flexible in cases where the tail behavior of data is important. It measures nonlinear dependence where the regular correlation…

Statistics Theory · Mathematics 2008-12-18 Zhengjun Zhang

In a standard cluster analysis, such as k-means, in addition to clusters locations and distances between them, it's important to know if they are connected or well separated from each other. The main focus of this paper is discovering the…

Machine Learning · Statistics 2017-05-22 Evgeny Bauman , Konstantin Bauman

In many situations, when dealing with several populations, equality of the covariance operators is assumed. An important issue is to study if this assumption holds before making other inferences. In this paper, we develop a test for…

Statistics Theory · Mathematics 2016-11-21 Graciela Boente , Daniela Rodriguez , Mariela Sued

Many data sources are naturally modeled by multiple weight assignments over a set of keys: snapshots of an evolving database at multiple points in time, measurements collected over multiple time periods, requests for resources served at…

Databases · Computer Science 2010-11-11 Edith Cohen , Haim Kaplan , Subhabrata Sen

In network data analysis, summary statistics of a network can provide us with meaningful insight into the structure of the network. The average clustering coefficient is one of the most popular and widely used network statistics. In this…

Statistics Theory · Mathematics 2023-11-21 Mingao Yuan , Xiaofeng Zhao

In observational studies of treatment effects, matched samples are created so treated and control groups are similar in terms of observable covariates. Traditionally such matched samples consist of matched pairs. If a pair match fails to…

Methodology · Statistics 2014-10-22 Luke Keele , Sam Pimentel , Frank Yoon

The average causal effect can often be best understood in the context of its variation. We demonstrate with two sets of four graphs, all of which represent the same average effect but with much different patterns of heterogeneity. As with…

Methodology · Statistics 2023-02-28 Andrew Gelman , Jessica Hullman , Lauren Kennedy

The coefficient of variation is a useful indicator for comparing the spread of values between dataset with different units or widely different means. In this paper we address the problem of investigating the equality of the coefficients of…

Methodology · Statistics 2023-06-06 Francesco Bertolino , Silvia Columbu , Mara Manca , Monica Musio

Copulas are mathematical objects that fully capture the dependence structure among random variables and hence, offer a great flexibility in building multivariate stochastic models. In statistics, a copula is used as a general way of…

Methodology · Statistics 2013-10-01 Abhik Ghosh , Aritra Chakravorty

A correlation is a binary vector that encodes all possible positions of overlaps of two words, where an overlap for an ordered pair of words (u,v) occurs if a suffix of word u matches a prefix of word v. As multiple pairs can have the same…

Discrete Mathematics · Computer Science 2025-06-03 Eric Rivals , Pengfei Wang

Clustered data are common in practice. Clustering arises when subjects are measured repeatedly, or subjects are nested in groups (e.g., households, schools). It is often of interest to evaluate the correlation between two variables with…

Methodology · Statistics 2025-01-16 Shengxin Tu , Chun Li , Bryan E. Shepherd

Comparison and contrast are the basic means to unveil causation and learn which treatments work. To build good comparison groups, randomized experimentation is key, yet often infeasible. In such non-experimental settings, we illustrate and…

Methodology · Statistics 2024-01-30 Ambarish Chattopadhyay , Jose R. Zubizarreta

In this paper, we present a new way of matching in observational studies that overcomes three limitations of existing matching approaches. First, it directly balances covariates with multi-valued treatments without requiring the generalized…

Applications · Statistics 2019-07-11 Magdalena Bennett , Juan Pablo Vielma , Jose R. Zubizarreta

We propose a summary measure defined as the expected value of a random variable over disjoint subsets of its support that are specified by a given grid of proportions, and consider its use in a regression modeling framework. The obtained…

Statistics Theory · Mathematics 2018-10-19 Celia García-Pareja , Matteo Bottai

Given a collection of computational models that all estimate values of the same natural process, we compare the performance of the average of the collection to the individual member whose estimates are nearest a given set of observations.…

Statistics Theory · Mathematics 2008-07-10 C. L. Winter , D. Nychka

Two-sample tests for multivariate data and especially for non-Euclidean data are not well explored. This paper presents a novel test statistic based on a similarity graph constructed on the pooled observations from the two samples. It can…

Methodology · Statistics 2024-08-12 Hao Chen , Jerome H. Friedman

This paper addresses how to calculate and interpret the time-delayed mutual information for a complex, diversely and sparsely measured, possibly non-stationary population of time-series of unknown composition and origin. The primary vehicle…

Chaotic Dynamics · Physics 2015-05-30 D. J. Albers , George Hripcsak

This work derives closed-form expressions computing the expectation of co-presence and of number of co-occurrences of nodes on paths sampled from a network according to general path weights (a bag of paths). The underlying idea is that two…

Machine Learning · Computer Science 2021-08-24 Guillaume Guex , Sylvain Courtain , Marco Saerens