中文
相关论文

相关论文: A note on data splitting with e-values: online app…

200 篇论文

We discuss systematically two versions of confidence regions: those based on p-values and those based on e-values, a recent alternative to p-values. Both versions can be applied to multiple hypothesis testing, and in this paper we are…

统计理论 · 数学 2024-03-05 Vladimir Vovk , Ruodu Wang

In this article we propose an optimal method referred to as SPlit for splitting a dataset into training and testing sets. SPlit is based on the method of Support Points (SP), which was initially developed for finding the optimal…

机器学习 · 统计学 2021-05-10 V. Roshan Joseph , Akhil Vakayil

Even though a train/test split of the dataset randomly performed is a common practice, could not always be the best approach for estimating performance generalization under some scenarios. The fact is that the usual machine learning…

机器学习 · 计算机科学 2022-09-09 Carlos Catania , Jorge Guerra , Juan Manuel Romero , Gabriel Caffaratti , Martin Marchetta

In this work, we develop a method named Twinning, for partitioning a dataset into statistically similar twin sets. Twinning is based on SPlit, a recently proposed model-independent method for optimally splitting a dataset into training and…

机器学习 · 统计学 2022-02-17 Akhil Vakayil , V. Roshan Joseph

Sampling is often a necessary evil to reduce the processing and storage costs of distributed tracing. In this work, we describe a scalable and adaptive sampling approach that can preserve events of interest better than the widely used…

数据结构与算法 · 计算机科学 2021-07-19 Otmar Ertl

E-values have gained prominence as flexible tools for statistical inference and risk control, enabling anytime- and post-hoc-valid procedures under minimal assumptions. However, many real-world applications fundamentally rely on sensitive…

统计方法学 · 统计学 2025-10-22 Daniel Csillag , Diego Mesquita

In recent years, an increasing amount of data is collected in different and often, not cooperative, databases. The problem of privacy-preserving, distributed calculations over separated databases and, a relative to it, issue of private data…

数据库 · 计算机科学 2016-05-23 Philip Derbeko , Shlomi Dolev , Ehud Gudes , Jeffrey D. Ullman

Bipartite Experiments are randomized experiments where the treatment is applied to a set of units (randomization units) that is different from the units of analysis, and randomization units and analysis units are connected through a…

统计方法学 · 统计学 2025-11-05 Liang Shi , Edvard Bakhitov , Kenneth Hung , Brian Karrer , Charlie Walker , Monica Bhole , Okke Schrijvers

The large-scale multiple testing inherent to high throughput biological data necessitates very high statistical stringency and thus true effects in data are difficult to detect unless they have high effect sizes. One solution to this…

统计方法学 · 统计学 2017-12-21 Mohamad S. Hasan

The most popular approach for analyzing survival data is the Cox regression model. The Cox model may, however, be misspecified, and its proportionality assumption may not always be fulfilled. An alternative approach for survival prediction…

机器学习 · 统计学 2018-05-17 Marvin N. Wright , Theresa Dankowski , Andreas Ziegler

We develop the theory of hypothesis testing based on the e-value, a notion of evidence that, unlike the p-value, allows for effortlessly combining results from several studies in the common scenario where the decision to perform a new study…

统计理论 · 数学 2023-03-13 Peter Grünwald , Rianne de Heide , Wouter Koolen

Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…

统计方法学 · 统计学 2024-09-05 F. Richard Guo , Rajen D. Shah

The p-values are often implicitly used as a measure of evidence for the hypotheses of the tests. This practice has been analyzed with different approaches. It is generally accepted for the one-sided hypothesis problem, but it is often…

统计理论 · 数学 2007-06-13 Guy Morel

In recent years, sparse principal component analysis has emerged as an extremely popular dimension reduction technique for high-dimensional data. The theoretical challenge, in the simplest case, is to estimate the leading eigenvector of a…

统计理论 · 数学 2016-09-29 Tengyao Wang , Quentin Berthet , Richard J. Samworth

Quality data is a fundamental contributor to success in statistics and machine learning. If a statistical assessment or machine learning leads to decisions that create value, data contributors may want a share of that value. This paper…

计算机科学与博弈论 · 计算机科学 2019-06-28 Eric Bax

We present the expected values from p-value hacking as a choice of the minimum p-value among $m$ independents tests, which can be considerably lower than the "true" p-value, even with a single trial, owing to the extreme skewness of the…

应用统计 · 统计学 2018-01-29 Nassim Nicholas Taleb

Effective methodologies for evaluating recommender systems are critical, so that such systems can be compared in a sound manner. A commonly overlooked aspect of recommender system evaluation is the selection of the data splitting strategy.…

信息检索 · 计算机科学 2020-07-28 Zaiqiao Meng , Richard McCreadie , Craig Macdonald , Iadh Ounis

The Cox model is an indispensable tool for time-to-event analysis, particularly in biomedical research. However, medicine is undergoing a profound transformation, generating data at an unprecedented scale, which opens new frontiers to study…

统计方法学 · 统计学 2023-03-07 Alexander W. Jung , Moritz Gerstung

Probabilistic graphical models have emerged as a powerful modeling tool for several real-world scenarios where one needs to reason under uncertainty. A graphical model's partition function is a central quantity of interest, and its…

人工智能 · 计算机科学 2021-05-25 Durgesh Agrawal , Yash Pote , Kuldeep S Meel

Sequential decision making significantly speeds up research and is more cost-effective compared to fixed-n methods. We present a method for sequential decision making for stratified count data that retains Type-I error guarantee or false…

统计方法学 · 统计学 2023-02-23 Rosanne J. Turner , Peter D. Grünwald