中文
相关论文

相关论文: The Anytime-Valid Logrank Test: Error Control Unde…

200 篇论文

The one-sample log-rank test is the method of choice for single-arm Phase II trials with time-to-event endpoint. It allows to compare the survival of the patients to a reference survival curve that typically represents the expected survival…

统计方法学 · 统计学 2026-03-02 Jannik Feld , Moritz Fabian Danzer , Andreas Faldum , Rene Schmidt

We propose a new algorithmic framework for sequential hypothesis testing with i.i.d. data, which includes A/B testing, nonparametric two-sample testing, and independence testing as special cases. It is novel in several ways: (a) it takes…

机器学习 · 统计学 2016-03-03 Akshay Balsubramani , Aaditya Ramdas

Nonparametric covariate adjustment is considered for log-rank type tests of treatment effect with right-censored time-to-event data from clinical trials applying covariate-adaptive randomization. Our proposed covariate-adjusted log-rank…

统计方法学 · 统计学 2023-01-23 Ting Ye , Jun Shao , Yanyao Yi

We introduce novel wild bootstrap procedures for testing superiority in unpaired two-sample survival data. By combining different classical weighted logrank test we obtain tests with broader power behavior. Right censoring within the data…

统计理论 · 数学 2018-08-20 Marc Ditzhaus , Markus Pauly

We propose a sequential, anytime-valid method to test the conditional independence of a response $Y$ and a predictor $X$ given a random vector $Z$. The proposed test is based on e-statistics and test martingales, which generalize likelihood…

统计方法学 · 统计学 2023-02-22 Peter Grünwald , Alexander Henzi , Tyron Lardy

E-variables are tools for retaining type-I error guarantee with optional stopping. We extend E-variables for sequential two-sample tests to general null hypotheses and anytime-valid confidence sequences. We provide implementations for…

统计方法学 · 统计学 2022-06-27 Rosanne Turner , Peter Grünwald

Time-to-event endpoints show an increasing popularity in phase II cancer trials. The standard statistical tool for such one-armed survival trials is the one-sample log-rank test. Its distributional properties are commonly derived in the…

统计方法学 · 统计学 2026-03-02 Moritz Fabian Danzer , Andreas Faldum , Rene Schmidt

We develop anytime-valid tests of invariance under the action of compact groups. The resulting test statistics are optimal in a logarithmic-growth sense. We apply our method to extend recent anytime-valid tests of independence and to…

统计方法学 · 统计学 2024-05-24 Tyron Lardy , Muriel Felipe Pérez-Ortiz

The proliferation of Large Language Models (LLMs) necessitates efficient mechanisms to distinguish machine-generated content from human text. While statistical watermarking has emerged as a promising solution, existing methods suffer from…

机器学习 · 计算机科学 2026-02-20 Baihe Huang , Eric Xu , Kannan Ramchandran , Jiantao Jiao , Michael I. Jordan

We consider the problem of independence testing for two univariate random variables in a sequential setting. By leveraging recent developments on safe, anytime-valid inference, we propose a test with time-uniform type I error control and…

统计方法学 · 统计学 2024-01-29 Alexander Henzi , Michael Law

Conditional Value-at-Risk (CVaR) is a central tail-risk measure in stochastic structural mechanics, yet its accurate evaluation under high-dimensional, spatially correlated material uncertainty remains computationally prohibitive for…

机器学习 · 统计学 2026-02-11 Alireza Tabarraei

In the past two decades, AB testing has proliferated to optimise products in digital domains. Traditional AB tests use fixed-horizon testing, determining the sample size of the experiment and continuing until the experiment has concluded.…

统计方法学 · 统计学 2023-11-01 Daniel Beasley

Iterative retrieval-augmented generation (RAG) enables large language models to answer complex multi-hop questions, but each additional loop increases latency, costs, and the risk of introducing distracting evidence, motivating the need for…

机器学习 · 计算机科学 2025-10-17 Jaewan Park , Solbee Cho , Jay-Yoon Lee

Widely used methods and software for group sequential tests of a null hypothesis of no treatment difference that allow for early stopping of a clinical trial depend primarily on the fact that sequentially-computed test statistics have the…

统计方法学 · 统计学 2025-06-19 Anastasios A. Tsiatis , Marie Davidian

The log-rank test and the Cox proportional hazards model are commonly used to compare time-to-event data in clinical trials, as they are most powerful under proportional hazards. But there is a loss of power if this assumption is violated,…

统计方法学 · 统计学 2024-02-14 Jonas Brugger , Tim Friede , Florian Klinglmüller , Martin Posch , Robin Ristl , Franz König

Studies to compare the survival of two or more groups using time-to-event data are of high importance in medical research. The gold standard is the log-rank test, which is optimal under proportional hazards. As the latter is no simple…

统计方法学 · 统计学 2022-10-25 Ina Dormuth , Tiantian Liu , Jin Xu , Markus Pauly , Marc Ditzhaus

Across many risk-sensitive areas, it is critical to continuously audit machine learning systems as we receive more data to quickly determine if they are performing as designed. This auditing task can be modeled as a sequential hypothesis…

机器学习 · 统计学 2026-05-26 Beepul Bharti , Ambar Pal , Jeremias Sulam

Lean is increasingly used to judge natural-language mathematical answers, but its signal is partial: many answers never formalize, and a failed proof may reflect an ill-typed statement or a missing library fact, not a wrong answer. On…

人工智能 · 计算机科学 2026-05-28 Pauline Bourigault , Xiaotong Ji , Matthieu Zimmer , Rasul Tutunov , Haitham Bou Ammar

We propose a new class of weighted logrank tests (WLRT) that control the risk of concluding that a new drug is more efficacious than standard of care, when, in fact, it is uniformly inferior. Perhaps surprisingly, this risk is not…

应用统计 · 统计学 2018-07-31 Dominic Magirr , Carl-Fredrik Burman

A/B tests are the gold standard for evaluating digital experiences on the web. However, traditional "fixed-horizon" statistical methods are often incompatible with the needs of modern industry practitioners as they do not permit continuous…

‹ 上一页 1 2 3 10 下一页 ›