English
Related papers

Related papers: Anytime-Valid Confidence Sequences in an Enterpris…

200 papers

Anytime valid sequential tests permit us to stop testing based on the current data, without invalidating the inference. Given a maximum number of observations $N$, one may believe this must come at the cost of power when compared to a…

Statistics Theory · Mathematics 2025-12-23 Nick W. Koning , Sam van Meer

Accelerated life tests (ALTs) play a crucial role in reliability analyses, providing lifetime estimates of highly reliable products. Among ALTs, step-stress design increases the stress level at predefined times, while maintaining a constant…

Statistics Theory · Mathematics 2024-02-12 Narayanaswamy Balakrishnan , María Jaenada , Leandro Pardo

A common problem in numerous research areas, particularly in clinical trials, is to test whether the effect of an explanatory variable on an outcome variable is equivalent across different groups. In practice, these tests are frequently…

Methodology · Statistics 2024-05-03 Niklas Hagemann , Kathrin Möllenhoff

Recently virtual platforms and virtual prototyping techniques have been widely applied for accelerating software development in electronics companies. It has been proved that these techniques can greatly shorten time-to-market and improve…

Software Engineering · Computer Science 2016-01-25 Bin Lin , Dejun Qian

Grounded text generation systems often generate text that contains factual inconsistencies, hindering their real-world applicability. Automatic factual consistency evaluation may help alleviate this limitation by accelerating evaluation…

Scenario-based testing has emerged as a common method for autonomous vehicles (AVs) safety assessment, offering a more efficient alternative to mile-based testing by focusing on high-risk scenarios. However, fundamental questions persist…

Software Engineering · Computer Science 2025-07-17 Xingyu Zhao , Robab Aghazadeh-Chakherlou , Chih-Hong Cheng , Peter Popov , Lorenzo Strigini

Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.…

Software Engineering · Computer Science 2026-05-13 Zhiqing Zhong , Zhijing Ye , Jiamin Wang , Xiaodong Yu

We propose a new algorithmic framework for sequential hypothesis testing with i.i.d. data, which includes A/B testing, nonparametric two-sample testing, and independence testing as special cases. It is novel in several ways: (a) it takes…

Machine Learning · Statistics 2016-03-03 Akshay Balsubramani , Aaditya Ramdas

Change point detection in high dimensional data has found considerable interest in recent years. Most of the literature either designs methodology for a retrospective analysis, where the whole sample is already available when the…

Statistics Theory · Mathematics 2020-12-16 Josua Gösmann , Christina Stoehr , Johannes Heiny , Holger Dette

Online experiments %in which experimental units receive a sequence of treatments over time are frequently employed in many technological companies to evaluate the performance of a newly developed policy, product, or treatment relative to a…

Econometrics · Economics 2025-01-14 Ke Sun , Linglong Kong , Hongtu Zhu , Chengchun Shi

Given that machine learning algorithms are increasingly being deployed to aid in high stakes decision-making, uncertainty quantification methods that wrap around these black box models such as conformal prediction have received much…

Machine Learning · Statistics 2026-02-09 Kayla E. Scharfstein , Arun Kumar Kuchibhotla

There have been major developments in Automated Driving (AD) and Driving Assist Systems (ADAS) in recent years. However, their safety assurance, thus methodologies for testing, verification and validation AD/ADAS safety-critical…

Software Engineering · Computer Science 2023-11-08 Hasan Esen , Brian Hsuan-Cheng Liao

Simulation can evaluate a statistical method for properties such as Type I Error, FDR, or bias on a grid of hypothesized parameter values. But what about the gaps between the grid-points? Continuous Simulation Extension (CSE) is a…

Methodology · Statistics 2024-09-10 James Yang , T. Ben Thompson , Michael Sklar

A/B testing refers to the statistical procedure of conducting an experiment to compare two treatments, A and B, applied to different testing subjects. It is widely used by technology companies such as Facebook, LinkedIn, and Netflix, to…

Methodology · Statistics 2026-05-12 Victoria Pokhiko , Qiong Zhang , Lulu Kang , D'arcy P. Mays

Quantum measurements under realistic conditions reveal only partial information about a system. Yet, by performing sequential measurements on the same system, additional information can be accessed. We investigate this problem in the…

Quantum Physics · Physics 2025-10-23 Carles Roch I Carceller , Hanwool Lee , Jonatan Bohr Brask , Kieran Flatt , Joonwoo Bae

The complexity of software in embedded systems has increased significantly over the last years so that software verification now plays an important role in ensuring the overall product quality. In this context, SAT-based bounded model…

Software Engineering · Computer Science 2009-11-20 Lucas Cordeiro , Bernd Fischer , Joao Marques-Silva

A/B testing remains the gold standard for evaluating modifications to e-commerce storefronts, yet it diverts traffic, requires weeks to reach statistical significance, and risks degrading user experience. We present SimGym, a framework for…

This paper presents ModelGuard, a sampling-based approach to runtime model validation for Lipschitz-continuous models. Although techniques exist for the validation of many classes of models the majority of these methods cannot be applied to…

Systems and Control · Electrical Eng. & Systems 2021-05-03 Taylor J. Carpenter , Radoslav Ivanov , Insup Lee , James Weimer

Online controlled experiments (also known as A/B Testing) have been viewed as a golden standard for large data-driven companies since the last few decades. The most common A/B testing framework adopted by many companies use "average…

Methodology · Statistics 2021-10-15 Yihan Bao , Shichao Han , Yong Wang

A standard practice in statistical hypothesis testing is to mention the p-value alongside the accept/reject decision. We show the advantages of mentioning an e-value instead. With p-values, it is not clear how to use an extreme observation…

Methodology · Statistics 2024-04-04 Peter Grünwald
‹ Prev 1 8 9 10 Next ›