中文
相关论文

相关论文: A note on data splitting with e-values: online app…

200 篇论文

This note is my comment on Glenn Shafer's discussion paper "Testing by betting", together with two online appendices comparing p-values and betting scores.

统计方法学 · 统计学 2021-05-07 Vladimir Vovk

Data splitting divides data into two parts. One part is reserved for model selection. In some applications, the second part is used for model validation but we use this part for estimating the parameters of the chosen model. We focus on the…

统计方法学 · 统计学 2016-01-20 Julian J. Faraway

Shafer (2021) offers a betting perspective on statistical testing which may be useful for foundational debates, given that disputes over such testing continue to be intense. To be helpful for researchers, however, this perspective will need…

统计理论 · 数学 2021-02-11 Sander Greenland

Multiple testing of a single hypothesis and testing multiple hypotheses are usually done in terms of p-values. In this paper we replace p-values with their natural competitor, e-values, which are closely related to betting, Bayes factors,…

统计理论 · 数学 2021-10-26 Vladimir Vovk , Ruodu Wang

Many statistical models require an estimation of unknown (co)-variance parameter(s) in a model. The estimation usually obtained by maximizing a log-likelihood which involves log determinant terms. In principle, one requires the…

统计计算 · 统计学 2016-09-05 Shengxin Zhu , Tongxiang Gu , Xiaowen Xu , Zeyao Mo

Conformal prediction is a powerful framework for distribution-free uncertainty quantification. The standard approach to conformal prediction relies on comparing the ranks of prediction scores: under exchangeability, the rank of a future…

机器学习 · 统计学 2025-05-07 Etienne Gauthier , Francis Bach , Michael I. Jordan

A recurring debate in the philosophy of statistics concerns what, exactly, should count as a measure of evidence for or against a given hypothesis. P-values, likelihood ratios, and Bayes factors all have their defenders. In this paper we…

统计方法学 · 统计学 2026-03-26 Ben Chugg , Aaditya Ramdas , Peter Grünwald

Probability forecasts for binary events play a central role in many applications. Their quality is commonly assessed with proper scoring rules, which assign forecasts a numerical score such that a correct forecast achieves a minimal…

统计方法学 · 统计学 2022-07-04 Alexander Henzi , Johanna F. Ziegel

We introduce a testing-by-betting framework that leverages predictions on unlabeled data to enhance the power of sequential hypothesis testing. Given limited samples from the joint distribution of $(X,Y)$, and additional unlabeled samples…

机器学习 · 计算机科学 2026-05-28 Yaniv Tenzer , Elad Tolochinsky , Yaniv Romano

We introduce a powerful deep classifier two-sample test for high-dimensional data based on E-values, called E-value Classifier Two-Sample Test (E-C2ST). Our test combines ideas from existing work on split likelihood ratio tests and…

统计方法学 · 统计学 2024-05-01 Teodora Pandeva , Tim Bakker , Christian A. Naesseth , Patrick Forré

We address the problem of testing conditional mean and conditional variance for non-stationary data. We build e-values and p-values for four types of non-parametric composite hypotheses with specified mean and variance as well as other…

统计理论 · 数学 2024-09-25 Yixuan Fan , Zhanyi Jiao , Ruodu Wang

A standard practice in statistical hypothesis testing is to mention the p-value alongside the accept/reject decision. We show the advantages of mentioning an e-value instead. With p-values, it is not clear how to use an extreme observation…

统计方法学 · 统计学 2024-04-04 Peter Grünwald

Assigning significance in high-dimensional regression is challenging. Most computationally efficient selection algorithms cannot guard against inclusion of noise variables. Asymptotically valid p-values are not available. An exception is a…

统计方法学 · 统计学 2009-06-12 Nicolai Meinshausen , Lukas Meier , Peter Bühlmann

The vast availability of large scale, massive and big data has increased the computational cost of data analysis. One such case is the computational cost of the univariate filtering which typically involves fitting many univariate…

统计方法学 · 统计学 2020-02-13 M. Tsagris , A. Alenazi , S. Fafalios

Hypothesis testing via e-variables can be framed as a sequential betting game, where a player each round picks an e-variable. A good player's strategy results in an effective statistical test that rejects the null hypothesis as soon as…

统计理论 · 数学 2025-05-30 Eugenio Clerico

Gorman and Bedrick (2019) argued for using random splits rather than standard splits in NLP experiments. We argue that random splits, like standard splits, lead to overly optimistic performance estimates. We can also split data in biased or…

计算与语言 · 计算机科学 2021-04-27 Anders Søgaard , Sebastian Ebert , Jasmijn Bastings , Katja Filippova

E-values have attracted considerable interest in recent years as flexible tools for enabling anytime-valid and adaptive data analysis. Hypothesis testing is at the core of many of these applications, which can often involve private or…

密码学与安全 · 计算机科学 2026-05-29 Ben Jacobsen , Tomas Gonzales , Gavin Brown , Kassem Fawaz , Aaditya Ramdas

Sample splitting is widely used in statistical applications, including classically in classification and more recently for inference post model selection. Motivating by problems in the study of diet, physical activity, and health, we…

统计方法学 · 统计学 2019-08-13 Eli S. Kravitz , Raymond J. Carroll , David Ruppert

We consider the problem of providing valid inference for a selected parameter in a sparse regression setting. It is well known that classical regression tools can be unreliable in this context due to the bias generated in the selection…

统计方法学 · 统计学 2022-12-07 Daniel G. Rasines , G. Alastair Young

We propose novel methodology for testing equality of model parameters between two high-dimensional populations. The technique is very general and applicable to a wide range of models. The method is based on sample splitting: the data is…

统计方法学 · 统计学 2013-01-17 Nicolas Städler , Sach Mukherjee
‹ 上一页 1 2 3 10 下一页 ›