Related papers: A modified $\chi^2$-test for uplift models with ap…
The ROC curve is the gold standard for measuring the performance of a test/scoring statistic regarding its capacity to discriminate between two statistical populations in a wide variety of applications, ranging from anomaly detection in…
We investigate the large-sample behavior of change-point tests based on weighted two-sample U-statistics, in the case of short-range dependent data. Under some mild mixing conditions, we establish convergence of the test statistic to an…
Statistical NLP systems are frequently evaluated and compared on the basis of their performances on a single split of training and test data. Results obtained using a single split are, however, subject to sampling noise. In this paper we…
As data shift or new data become available, updating clinical machine learning models may be necessary to maintain or improve performance over time. However, updating a model can introduce compatibility issues when the behavior of the…
Models that top leaderboards often perform unsatisfactorily when deployed in real world applications; this has necessitated rigorous and expensive pre-deployment model testing. A hitherto unexplored facet of model performance is: Are our…
We study a tractable two-player contest built on a truncated cubic contest success function. Its defining feature is a strategic-feedback parameter whose sign determines whether a leading player's effort lowers (suppression) or raises…
In this paper, four new Chi-Square type statistics are presented for testing the hypothesis of a uniform null versus specified trend alternatives. The powers of these test statistics are compared with the powers of the statistics considered…
Experimental comparisons of performance represent an important aspect of research on optimization algorithms. In this work we present a methodology for defining the required sample sizes for designing experiments with desired statistical…
Surveys are commonly used to facilitate research in epidemiology, health, and the social and behavioral sciences. Often, these surveys are not simple random samples, and respondents are given weights reflecting their probability of…
This study replicates and adapts the experiment of Hoelzl and Rustichini (2005), which examined overplacement, i.e., overconfidence in relative self-assessments, by analyzing individuals' voting preferences between a performance-based and a…
The causal effect of showing an ad to a potential customer versus not, commonly referred to as "incrementality", is the fundamental question of advertising effectiveness. In digital advertising three major puzzle pieces are central to…
The test statistics of two powerful tests for normality \citep{lm1,mud2} are estimators of the correlation coefficient between certain sample moments. We derive new versions of the test statistics that are functions of the sample skewness…
The raking-ratio method is a statistical and computational method which adjusts the empirical measure to match the true probability of sets of a finite partition. We study the asymptotic behavior of the raking-ratio empirical process…
We consider the conditional randomization test as a way to account for covariate imbalance in randomized experiments. The test accounts for covariate imbalance by comparing the observed test statistic to the null distribution of the test…
We introduce a new conservative test for quantifying the consistency of two or more datasets. The test is based on the Bayesian answer to the question, ``How much more probable is it that all my data were generated from the same model…
Estimating the strength of dependency between two variables is fundamental for exploratory analysis and many other applications in data mining. For example: non-linear dependencies between two continuous variables can be explored with the…
This paper introduces a marketing decision framework that optimizes customer targeting by integrating heterogeneous treatment effect estimation with explicit business guardrails. The objective is to maximize revenue and retention while…
The median statistic has recently been discussed by Gott \textit{et al.} as a more reliable alternative to the standard $\chi^2$ likelihood analysis, in the sense of requiring fewer assumptions about the data and being almost as…
We consider whether the asymptotic distributions for the log-likelihood ratio test statistic are expected to be Gaussian or chi-squared. Two straightforward examples provide insight on the difference.
The odds ratio measure is used in health and social surveys where the odds of a certain event is to be compared between two populations. It is defined using logistic regression, and requires that data from surveys are accompanied by their…