English
Related papers

Related papers: $t$-Testing the Waters: Empirically Validating Ass…

200 papers

Online controlled experiments (a.k.a. A/B testing) have been used as the mantra for data-driven decision making on feature changing and product shipping in many Internet companies. However, it is still a great challenge to systematically…

Applications · Statistics 2018-08-16 Yuxiang Xie , Nanyu Chen , Xiaolin Shi

The parametric Welch $t$-test and the non-parametric Wilcoxon-Mann-Whitney test are the most commonly used two independent sample means tests. More recent testing approaches include the non-parametric, empirical likelihood and exponential…

Methodology · Statistics 2019-10-08 Michail Tsagris , Abdulaziz Alenazi , Kleio-Maria Verrou , Nikolaos Pandis

Experimentation platforms in industry must often deal with customer trust issues. Platforms must prove the validity of their claims as well as catch issues that arise. As a central quantity estimated by experimentation platforms, the…

Methodology · Statistics 2025-11-21 Kedar Karhadkar , Jack Klys , Daniel Ting , Artem Vorozhtsov , Houssam Nassif

Hotelling's T-squared test is a classical tool to test if the normal mean of a multivariate normal distribution is a specified one or the means of two multivariate normal means are equal. When the population dimension is higher than the…

Statistics Theory · Mathematics 2021-08-17 Tiefeng Jiang , Ping Li

In this article, we try to give an answer to the simple question: ``\textit{What is the critical growth rate of the dimension $p$ as a function of the sample size $n$ for which the Central Limit Theorem holds uniformly over the collection…

Probability · Mathematics 2020-08-12 Debraj Das , Soumendra Lahiri

Current approaches to A/B testing in networks focus on limiting interference, the concern that treatment effects can "spill over" from treatment nodes to control nodes and lead to biased causal effect estimation. Prominent methods for…

Machine Learning · Computer Science 2020-04-16 Zahra Fatemi , Elena Zheleva

Randomized Controlled Trials (RCT)s are relied upon to assess new treatments, but suffer from limited power to guide personalized treatment decisions. On the other hand, observational (i.e., non-experimental) studies have large and diverse…

Methodology · Statistics 2023-03-07 Zeshan Hussain , Ming-Chieh Shih , Michael Oberst , Ilker Demirel , David Sontag

Cognitive diagnosis models have been popularly used in fields such as education, psychology, and social sciences. While parametric likelihood estimation is a prevailing method for fitting cognitive diagnosis models, nonparametric…

Statistics Theory · Mathematics 2025-10-01 Chengyu Cui , Yanlong Liu , Gongjun Xu

Experiments often yield non-identically distributed data for statistical analysis. Tests of hypothesis under such set-ups are generally performed using the likelihood ratio test, which is non-robust with respect to outliers and model…

Statistics Theory · Mathematics 2017-07-25 Abhik Ghosh , Ayanendranath Basu

A lot of online marketing campaigns aim to promote user interaction. The average treatment effect (ATE) of campaign strategies need to be monitored throughout the campaign. A/B testing is usually conducted for such needs, whereas the…

Social and Information Networks · Computer Science 2023-04-28 Tianchi Cai , Daxi Cheng , Chen Liang , Ziqi Liu , Lihong Gu , Huizhi Xie , Zhiqiang Zhang , Xiaodong Zeng , Jinjie Gu

We study user sentiment (reported via optional surveys) as a metric for fully randomized A/B tests. Both user-level covariates and treatment assignment can impact response propensity. We propose a set of consistent estimators for the…

Methodology · Statistics 2019-06-27 Ercan Yildiz , Joshua Safyan , Marc Harper

The law of large numbers (LLN) and central limit theorem (CLT) are long and widely been known as two fundamental results in probability theory. Recently problems of model uncertainties in statistics, measures of risk and superhedging in…

Probability · Mathematics 2007-05-23 Shige Peng

We propose a new framework for online testing of heterogeneous treatment effects. The proposed test, named sequential score test (SST), is able to control type I error under continuous monitoring and detect multi-dimensional heterogeneous…

Methodology · Statistics 2020-02-11 Miao Yu , Wenbin Lu , Rui Song

In the big data era, the need to reevaluate traditional statistical methods is paramount due to the challenges posed by vast datasets. While larger samples theoretically enhance accuracy and hypothesis testing power without increasing false…

Methodology · Statistics 2026-01-09 Xuekui Zhang , Li Xing , Jing Zhang , Soojeong Kim

The lack of non-parametric statistical tests for confounding bias significantly hampers the development of robust, valid and generalizable predictive models in many fields of research. Here I propose the partial and full confounder tests,…

Machine Learning · Computer Science 2025-05-30 Tamas Spisak

This paper proposes a new deep-learning method to construct test statistics by computer vision and metrics learning. The application highlighted in this paper is applying computer vision on Q-Q plot to construct a new test statistic for…

Computer Vision and Pattern Recognition · Computer Science 2019-09-17 Ke-Wei Huang , Mengke Qiao , Xuanqi Liu , Siyuan Liu , Mingxi Dai

In this paper we present randomization methods to enhance the accuracy of the central limit theorem (CLT) based inferences about the population mean $\mu$. We introduce a broad class of randomized versions of the Student $t$-statistic, the…

Methodology · Statistics 2016-05-20 Masoud M Nasari

Combining test statistics from independent trials or experiments is a popular method of meta-analysis. However, there is very limited theoretical understanding of the power of the combined test, especially in high-dimensional models…

Statistics Theory · Mathematics 2023-10-31 Botond Szabó , Aad van der Vaart , Lasse Vuursteen , Harry van Zanten

P values or risk ratios from multiple, independent studies, observational or randomized, can be computationally combined to provide an overall assessment of a research question in meta-analysis. There is a need to examine the reliability of…

Methodology · Statistics 2021-10-28 S. Stanley Young , Warren B. Kindzierski

What can be considered an appropriate statistical method for the primary analysis of a randomized clinical trial (RCT) with a time-to-event endpoint when we anticipate non-proportional hazards owing to a delayed effect? This question has…

Methodology · Statistics 2023-04-18 José L. Jiménez , Isobel Barrott , Francesca Gasperoni , Dominic Magirr