English
Related papers

Related papers: Online Experimentation with Surrogate Metrics: Gui…

200 papers

Linear models are foundational tools in statistics and ubiquitous across the applied sciences. However, conventional statistical inference -- such as $t$-tests and $F$-tests -- are only valid at fixed sample sizes, making them unsuitable…

Methodology · Statistics 2025-07-08 Michael Lindon , Dae Woong Ham , Martin Tingley , Iavor Bojinov

A common practice in clinical trials is to evaluate a treatment effect on an intermediate endpoint when the true outcome of interest would be difficult or costly to measure. We consider how to validate intermediate endpoints in a…

Methodology · Statistics 2022-11-30 Emily K. Roberts , Michael R. Elliott , Jeremy M. G. Taylor

Motivated by increasing pressure for decision makers to shorten the time required to evaluate the efficacy of a treatment such that treatments deemed safe and effective can be made publicly available, there has been substantial recent…

Methodology · Statistics 2022-09-20 Xuan Wang , Layla Parast , Lu Tian , Tianxi Cai

This paper examines how spillover effects in A/B testing can impede organizational progress and develops strategies for mitigating these challenges. We identify a phenomenon termed ``seesaw experimentation'', where a firm's overall…

General Economics · Economics 2025-01-16 Jin Li , Ye Luo , Xiaowei Zhang

We consider optimization of generalized performance metrics for binary classification by means of surrogate losses. We focus on a class of metrics, which are linear-fractional functions of the false positive and false negative rates…

Machine Learning · Computer Science 2016-10-10 Wojciech Kotłowski , Krzysztof Dembczyński

Surrogate neural network-based models have been lately trained and used in a variety of science and engineering applications where the number of evaluations of a target function is limited by execution time. In cell phone camera systems,…

Computational Engineering, Finance, and Science · Computer Science 2022-06-29 Shantanu Shahane , Erman Guleryuz , Diab W Abueidda , Allen Lee , Joe Liu , Xin Yu , Raymond Chiu , Seid Koric , Narayana R Aluru , Placid M Ferreira

Design of experiments and estimation of treatment effects in large-scale networks, in the presence of strong interference, is a challenging and important problem. Most existing methods' performance deteriorates as the density of the network…

Methodology · Statistics 2020-12-15 Preetam Nandy , Kinjal Basu , Shaunak Chatterjee , Ye Tu

Offline optimization is an important task in numerous material engineering domains where online experimentation to collect data is too expensive and needs to be replaced by an in silico maximization of a surrogate of the black-box function.…

Machine Learning · Computer Science 2025-03-07 Manh Cuong Dao , Phi Le Nguyen , Thao Nguyen Truong , Trong Nghia Hoang

Many real-world problems have expensive-to-compute fitness functions and are multi-objective in nature. Surrogate-assisted evolutionary algorithms are often used to tackle such problems. Despite this, literature about analysing the fitness…

Neural and Evolutionary Computing · Computer Science 2024-04-11 C. J. Rodriguez , S. L. Thomson , T. Alderliesten , P. A. N. Bosman

Online controlled experiments (A/B tests) are fundamental to data-driven decision-making in the digital economy. However, their real-world application is frequently compromised by two critical shortcomings: the use of statistically flawed…

Applications · Statistics 2025-09-30 Srijesh Pillai , Rajesh Kumar Chandrawat

Machine learning models of accelerator systems (`surrogate models') are able to provide fast, accurate predictions of accelerator physics phenomena. However, approaches to date typically do not include measured input diagnostics, such as…

Accelerator Physics · Physics 2021-04-06 Lipi Gupta , Auralee Edelen , Nicole Neveu , Aashwin Mishra , Christopher Mayes , Young-Kee Kim

Surrogate testing is used widely to determine the nature of the process generating the given empirical sample. In the present study, the usefulness of phase-randomized surrogates, amplitude adjusted Fourier transform (AAFT) and iterated…

Statistical Mechanics · Physics 2007-05-23 Radhakrishnan Nagarajan

Failure to accurately measure the outcomes of an experiment can lead to bias and incorrect conclusions. Online controlled experiments (aka AB tests) are increasingly being used to make decisions to improve websites as well as mobile and…

Other Computer Science · Computer Science 2019-04-01 Jayant Gupchup , Yasaman Hosseinkashi , Pavel Dmitriev , Daniel Schneider , Ross Cutler , Andrei Jefremov , Martin Ellis

The first part of this thesis focuses on maximizing the overall recommendation accuracy. This accuracy is usually evaluated with some user-oriented metric tailored to the recommendation scenario, but because recommendation is usually…

Information Retrieval · Computer Science 2023-11-14 Roger Zhe Li

It is increasingly common in digital environments to use A/B tests to compare the performance of recommendation algorithms. However, such experiments often violate the stable unit treatment value assumption (SUTVA), particularly SUTVA's "no…

In streaming platforms churn is extremely costly, yet A/B tests are typically evaluated using outcomes observed within a limited experimental horizon. Even when both short- and predicted long-term engagement metrics are considered, they may…

Machine Learning · Computer Science 2026-04-23 Dario Simionato , Andrea Tonon , Mingxue Wang , Weiguo Wang , Tong Gui , Xiaoyue Li

A/B tests serve the purpose of reliably identifying the effect of changes introduced in online services. It is common for online platforms to run a large number of simultaneous experiments by splitting incoming user traffic randomly in…

Machine Learning · Computer Science 2022-10-18 Alexander Buchholz , Vito Bellini , Giuseppe Di Benedetto , Yannik Stein , Matteo Ruffini , Fabian Moerchen

Online controlled experiment (also called A/B test or experiment) is the most important tool for decision-making at a wide range of data-driven companies like Microsoft, Google, Meta, etc. Metric computation is the core procedure for…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-09-27 Tao Xiong , Yong Wang

Empirical researchers often trim observations with small denominator A when they estimate moments of the form E[B/A]. Large trimming is a common practice to mitigate variance, but it incurs large trimming bias. This paper provides a novel…

Methodology · Statistics 2021-01-12 Yuya Sasaki , Takuya Ura

The identification of surrogate markers is motivated by their potential to make decisions sooner about a treatment effect. However, few methods have been developed to actually use a surrogate marker to test for a treatment effect in a…

Methodology · Statistics 2024-09-17 Layla Parast , Jay Bartroff