中文
相关论文

相关论文: Learning Metrics that Maximise Power for Accelerat…

200 篇论文

Two-sided marketplaces are standard business models of many online platforms (e.g., Amazon, Facebook, LinkedIn), wherein the platforms have consumers, buyers or content viewers on one side and producers, sellers or content-creators on the…

社会与信息网络 · 计算机科学 2021-10-28 Preetam Nandy , Divya Venugopalan , Chun Lo , Shaunak Chatterjee

Time series anomaly detection is widely used in IoT and cyber-physical systems, yet its evaluation remains challenging due to diverse application objectives and heterogeneous metric assumptions. This study introduces a problem-oriented…

人工智能 · 计算机科学 2026-05-15 Kaixiang Yang , Jiarong Liu , Yupeng Song , Shuanghua Yang , Yujue Zhou

Machine learning models excel with abundant annotated data, but annotation is often costly and time-intensive. Active learning (AL) aims to improve the performance-to-annotation ratio by using query methods (QMs) to iteratively select the…

机器学习 · 计算机科学 2026-02-17 Hannes Kath , Thiago S. Gouvêa , Daniel Sonntag

Science students must deal with the errors inherent to all physical measurements and be conscious of the need to expressvthem as a best estimate and a range of uncertainty. Errors are routinely classified as statistical or systematic.…

物理教育 · 物理学 2021-05-05 Martin Monteiro , Cecilia Stari , Cecilia Cabeza , Arturo C. Marti

Adaptive online testing efficiently assesses examinee proficiency by dynamically adjusting the difficulty of test items based on their performance. To achieve this, items are selected so that their difficulty closely matches the test…

统计方法学 · 统计学 2025-11-21 Hideo Hirose

Real-world recommender systems often need to balance multiple objectives when deciding which recommendations to present to users. These include behavioural signals (e.g. clicks, shares, dwell time), as well as broader objectives (e.g.…

信息检索 · 计算机科学 2024-09-17 Olivier Jeunen , Jatin Mandav , Ivan Potapov , Nakul Agarwal , Sourabh Vaid , Wenzhe Shi , Aleksei Ustimenko

We develop a theoretical framework for sample splitting in A/B testing environments, where data for each test are partitioned into two splits to measure methodological performance when the true impacts of tests are unobserved. We show that…

计量经济学 · 经济学 2026-03-24 Ryan Kessler , James McQueen , Miikka Rokkanen

Many large-scale testing procedures learn signal structure from the data to boost power. Direct data reuse can inflate Type-I error ("double dipping"), so a common remedy is masking: withholding some information during learning and using it…

统计理论 · 数学 2026-04-02 Abhinav Chakraborty , Junu Lee , Eugene Katsevich

Online experimentation, also known as A/B testing, is the gold standard for measuring product impacts and making business decisions in the tech industry. The validity and utility of experiments, however, hinge on unbiasedness and sufficient…

应用统计 · 统计学 2020-12-17 Min Liu , Jialiang Mao , Kang Kang

Sequential Multiple-Assignment Randomized Trials (SMARTs) play an increasingly important role in psychological and behavioral health research. This experimental approach enables researchers to answer scientific questions about how to…

统计方法学 · 统计学 2023-06-21 John J. Dziak , Daniel Almirall , Walter Dempsey , Catherine Stanger , Inbal Nahum-Shani

Companies offering web services routinely run randomized online experiments to estimate the causal impact associated with the adoption of new features and policies on key performance metrics of interest. These experiments are used to…

统计方法学 · 统计学 2023-07-13 Lorenzo Masoero , Doug Hains , James McQueen

A/B-tests are a cornerstone of experimental design on the web, with wide-ranging applications and use-cases. The statistical $t$-test comparing differences in means is the most commonly used method for assessing treatment effects, often…

统计方法学 · 统计学 2025-02-25 Olivier Jeunen

A/B tests are the gold standard for evaluating digital experiences on the web. However, traditional "fixed-horizon" statistical methods are often incompatible with the needs of modern industry practitioners as they do not permit continuous…

Data-driven most powerful tests are statistical hypothesis decision-making tools that deliver the greatest power against a fixed null hypothesis among all corresponding data-based tests of a given size. When the underlying data…

统计理论 · 数学 2023-03-15 Albert Vexler , Alan D. Hutson

Linear models are foundational tools in statistics and ubiquitous across the applied sciences. However, conventional statistical inference -- such as $t$-tests and $F$-tests -- are only valid at fixed sample sizes, making them unsuitable…

统计方法学 · 统计学 2025-07-08 Michael Lindon , Dae Woong Ham , Martin Tingley , Iavor Bojinov

Modern online experimentation faces two bottlenecks: scarce traffic forces tough choices on which variants to test, and post-hoc insight extraction is manual, inconsistent, and often content-agnostic. Meanwhile, organizations underuse…

人工智能 · 计算机科学 2026-02-17 Zhengmian Hu , Lei Shi , Ritwik Sinha , Justin Grover , David Arbour

To effectively test parts of the Internet of Things (IoT) systems with a state machine character, Model-based Testing (MBT) approach can be taken. In MBT, a system model is created, and test cases are generated automatically from the model,…

软件工程 · 计算机科学 2020-05-21 Vaclav Rechtberger , Miroslav Bures , Bestoun S. Ahmed

The goal of this article is to investigate how human participants allocate their limited time to decisions with different properties. We report the results of two behavioral experiments. In each trial of the experiments, the participant…

神经元与认知 · 定量生物学 2016-07-20 Arash Khodadadi , Pegah Fakhari , Jerome R. Busemeyer

Bandit algorithms are widely used in sequential decision problems to maximize the cumulative reward. One potential application is mobile health, where the goal is to promote the user's health through personalized interventions based on user…

机器学习 · 统计学 2022-08-23 Gi-Soo Kim , Hyun-Joon Yang , Jane P. Kim

Failure to accurately measure the outcomes of an experiment can lead to bias and incorrect conclusions. Online controlled experiments (aka AB tests) are increasingly being used to make decisions to improve websites as well as mobile and…

其他计算机科学 · 计算机科学 2019-04-01 Jayant Gupchup , Yasaman Hosseinkashi , Pavel Dmitriev , Daniel Schneider , Ross Cutler , Andrei Jefremov , Martin Ellis