中文
相关论文

相关论文: $t$-Testing the Waters: Empirically Validating Ass…

200 篇论文

Classical tests of fit typically reject a model for large enough real data samples. In contrast, often in statistical practice a model offers a good description of the data even though it is not the "true" random generator. We consider a…

统计理论 · 数学 2019-11-22 Eustasio del Barrio , Hristo Inouzhe , Carlos Matrán

Online controlled experiments are a crucial tool to allow for confident decision-making in technology companies. A North Star metric is defined (such as long-term revenue or user retention), and system variants that statistically…

机器学习 · 计算机科学 2024-06-14 Olivier Jeunen , Aleksei Ustimenko

Concerns have been expressed over the validity of statistical inference under covariate-adaptive randomization despite the extensive use in clinical trials. In the literature, the inferential properties under covariate-adaptive…

统计方法学 · 统计学 2022-07-05 Li Yang , Wei Ma , Yichen Qin , Feifang Hu

With the extensive use of digital devices, online experimental platforms are commonly used to conduct experiments to collect data for evaluating different variations of products, algorithms, and interface designs, a.k.a., A/B tests. In…

统计方法学 · 统计学 2024-07-09 Qiong Zhang , Lulu Kang , Xinwei Deng

A/B testing is a standard approach for evaluating the effect of online experiments; the goal is to estimate the `average treatment effect' of a new feature or condition by exposing a sample of the overall population to it. A drawback with…

社会与信息网络 · 计算机科学 2013-05-31 Johan Ugander , Brian Karrer , Lars Backstrom , Jon Kleinberg

Online A/B tests have become increasingly popular and important for social platforms. However, accurately estimating the global average treatment effect (GATE) has proven to be challenging due to network interference, which violates the…

统计方法学 · 统计学 2023-11-27 Qianyi Chen , Bo Li , Lu Deng , Yong Wang

Experimental testing is vital in the optimization of web applications, and as such A/B testing has been widely adopted as a methodology for determining optimal content for many web applications. While some testing platforms provide…

统计方法学 · 统计学 2017-10-04 Ian E. Fellows

Let $\mathbf{A}=\frac{1}{\sqrt{np}}(\mathbf{X}^T\mathbf{X}-p\mathbf {I}_n)$ where $\mathbf{X}$ is a $p\times n$ matrix, consisting of independent and identically distributed (i.i.d.) real random variables $X_{ij}$ with mean zero and…

统计理论 · 数学 2015-06-02 Binbin Chen , Guangming Pan

Assessment of multimedia quality relies heavily on subjective assessment, and is typically done by human subjects in the form of preferences or continuous ratings. Such data is crucial for analysis of different multimedia processing…

多媒体 · 计算机科学 2018-01-26 Manish Narwaria , Lukas Krasula , Patrick Le Callet

Rigorous statistical evaluations of large language models (LLMs), including valid error bars and significance testing, are essential for meaningful and reliable performance assessment. Currently, when such statistical measures are reported,…

人工智能 · 计算机科学 2025-05-29 Sam Bowyer , Laurence Aitchison , Desi R. Ivanova

In industry, online randomized controlled experiment (a.k.a. A/B experiment) is a standard approach to measure the impact of a causal change. These experiments have small treatment effect to reduce the potential blast radius. As a result,…

计量经济学 · 经济学 2025-05-29 Tanmoy Das , Dohyeon Lee , Arnab Sinha

A key feature of a sequential study is that the actual sample size is a random variable that typically depends on the outcomes collected. While hypothesis testing theory for sequential designs is well established, parameter and precision…

统计理论 · 数学 2017-12-21 Ben Berckmoes , Geert Molenberghs

Online experiments %in which experimental units receive a sequence of treatments over time are frequently employed in many technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a…

计量经济学 · 经济学 2025-01-14 Ke Sun , Linglong Kong , Hongtu Zhu , Chengchun Shi

A/B testing, a widely used form of Randomized Controlled Trial (RCT), is a fundamental tool in business data analysis and experimental design. However, despite its intent to maintain randomness, A/B testing often faces challenges that…

统计方法学 · 统计学 2024-08-13 Zihao Zheng , Carol Liu

A/B tests serve the purpose of reliably identifying the effect of changes introduced in online services. It is common for online platforms to run a large number of simultaneous experiments by splitting incoming user traffic randomly in…

While randomized trials may be the gold standard for evaluating the effectiveness of the treatment intervention, in some special circumstances, single-arm clinical trials utilizing external control may be considered. The causal treatment…

统计方法学 · 统计学 2025-05-26 Huan Wang , Fei Wu , Yeh-Fong Chen

Estimating the maximum mean finds a variety of applications in practice. In this paper, we study estimation of the maximum mean using an upper confidence bound (UCB) approach where the sampling budget is adaptively allocated to one of the…

统计理论 · 数学 2024-08-09 Zhang Kun , Liu Guangwu , Shi Wen

We develop a central limit theorem (CLT) for a non-parametric estimator of the transition matrices in controlled Markov chains (CMCs) with finite state-action spaces. Our results establish precise conditions on the logging policy under…

统计理论 · 数学 2026-03-26 Ziwei Su , Imon Banerjee , Diego Klabjan

Statistical techniques are used in all branches of science to determine the feasibility of quantitative hypotheses. One of the most basic applications of statistical techniques in comparative analysis is the test of equality of two…

统计方法学 · 统计学 2018-05-01 Ayanendranath Basu , Abhijit Mandal , Nirian Martin , Leandro Pardo

Online experimentation (or A/B testing) has been widely adopted in industry as the gold standard for measuring product impacts. Despite the wide adoption, few literatures discuss A/B testing with quantile metrics. Quantile metrics, such as…

应用统计 · 统计学 2019-03-22 Min Liu , Xiaohui Sun , Maneesh Varshney , Ya Xu