English
Related papers

Related papers: It's About Time: What A/B Test Metrics Estimate

200 papers

Online experiments (A/B tests) are widely regarded as the gold standard for evaluating recommender system variants and guiding launch decisions. However, a variety of biases can distort the results of the experiment and mislead…

Information Retrieval · Computer Science 2025-09-03 Chen Zheng , Zhenyu Zhao

In this paper we investigate the question of how much combined measurements can increase the accuracy of additive quantities. Therefore, we consider a set of measurements from a selection of all possible combinations of the $n$ labeled…

Data Analysis, Statistics and Probability · Physics 2022-05-18 B. Mirbach , M. Boguslawski

Social, also called human-aware, navigation is a key challenge for the integration of mobile robots into human environments. The evaluation of such systems is complex, as factors such as comfort, safety, and legibility must be considered.…

In many modern settings, data are acquired iteratively over time, rather than all at once. Such settings are known as online, as opposed to offline or batch. We introduce a simple technique for online parameter estimation, which can operate…

Computation · Statistics 2017-03-22 Hien D Nguyen

Very little attention has been paid to the comparison of efficiency between high accuracy statistical parsers. This paper proposes one machine-independent metric that is general enough to allow comparisons across very different parsing…

Computation and Language · Computer Science 2007-05-23 Brian Roark , Eugene Charniak

Mutual information is a general statistical dependency measure which has found applications in representation learning, causality, domain generalization and computational biology. However, mutual information estimators are typically…

Machine Learning · Statistics 2023-10-17 Paweł Czyż , Frederic Grabowski , Julia E. Vogt , Niko Beerenwinkel , Alexander Marx

Adaptive experiments are used extensively in online platforms, healthcare and biotechnology, and a variety of other settings. In many of these applications, the main goal is not to precisely estimate a treatment effect, but to demonstrate…

Statistics Theory · Mathematics 2026-03-10 Guido Imbens , Lorenzo Masoero , Alexander Rakhlin , Thomas S. Richardson , Suhas Vijaykumar

Survey respondents may give untruthful answers to sensitive questions when asked directly. In recent years, researchers have turned to the list experiment (also known as the item count technique) to overcome this difficulty. While list…

Applications · Statistics 2014-06-03 Peter M. Aronow , Alexander Coppock , Forrest W. Crawford , Donald P. Green

With the passage of more time from the original date of publication, the measure of the impact of scientific works using subsequent citation counts becomes more accurate. However the measurement of individual and organizational research…

Digital Libraries · Computer Science 2018-11-06 Giovanni Abramo , Tindaro Cicero , Ciriaco Andrea D'Angelo

Background: Measurement errors in terms of quantification or classification frequently occur in epidemiologic data and can strongly impact inference. Measurement errors may occur when ascertaining, recording or extracting data. Although the…

Methodology · Statistics 2021-10-22 Walter K Kremers

Email communication between instructors and students is ubiquitous, and it could be valuable to explore ways of testing out how to make email messages more impactful. This paper explores the design space of using emails to get students to…

Human-Computer Interaction · Computer Science 2022-08-11 Angela Zavaleta-Bernuy , Ziwen Han , Hammad Shaikh , Qi Yin Zheng , Lisa-Angelique Lim , Anna Rafferty , Andrew Petersen , Joseph Jay Williams

Measuring performance & quantifying a performance change are core evaluation techniques in programming language and systems research. Of 122 recent scientific papers, as many as 65 included experimental evaluation that quantified a…

Methodology · Statistics 2020-07-22 Tomas Kalibera , Richard Jones

Predicting when an individual will adopt a new behavior is an important problem in application domains such as marketing and public health. This paper examines the perfor- mance of a wide variety of social network based measurements…

Social and Information Networks · Computer Science 2016-07-26 Nikhil Kumar , Ruocheng Guo , Ashkan Aleali , Paulo Shakarian

Typically, machine learning models are trained and evaluated without making any distinction between users (e.g, using traditional hold-out and cross-validation). However, this produces inaccurate performance metrics estimates in multi-user…

Machine Learning · Computer Science 2023-12-11 Enrique Garcia-Ceja , Luciano Garcia-Banuelos , Nicolas Jourdan

Online controlled experimentation is widely adopted for evaluating new features in the rapid development cycle for web products and mobile applications. Measurement of the overall experiment sample is a common practice to quantify the…

Human-Computer Interaction · Computer Science 2022-01-27 Zhenyu Zhao , Yan He , Miao Chen

Online and AI-based symptom checkers are applications that assist medical laypeople in diagnosing their symptoms and determining which course of action to take. When evaluating these tools, previous studies primarily used an approach…

Human-Computer Interaction · Computer Science 2025-06-30 Marvin Kopka , Markus A. Feufel

There is growing interest in Bayesian clinical trial designs with informative prior distributions, e.g. for extrapolation of adult data to pediatrics, or use of external controls. While the classical type I error is commonly used to…

Methodology · Statistics 2023-09-06 Nicky Best , Maxine Ajimi , Beat Neuenschwander , Gaelle Saint-Hilary , Simon Wandel

The selection of the assumed effect size (AES) critically determines the duration of an experiment, and hence its accuracy and efficiency. Traditionally, experimenters determine AES based on domain knowledge. However, this method becomes…

Machine Learning · Computer Science 2025-04-15 Yu Liu , Runzhe Wan , James McQueen , Doug Hains , Jinxiang Gu , Rui Song

Recommender systems have become an integral part of online platforms, providing personalized recommendations for purchases, content consumption, and interpersonal connections. These systems consist of two sides: the producer side comprises…

Methodology · Statistics 2023-11-08 Yan Wang , Shan Ba

A/B testing, a widely used form of Randomized Controlled Trial (RCT), is a fundamental tool in business data analysis and experimental design. However, despite its intent to maintain randomness, A/B testing often faces challenges that…

Methodology · Statistics 2024-08-13 Zihao Zheng , Carol Liu
‹ Prev 1 8 9 10 Next ›