English
Related papers

Related papers: A note on data splitting with e-values: online app…

200 papers

We propose a new method, probabilistic divide-and-conquer, for improving the success probability in rejection sampling. For the example of integer partitions, there is an ideal recursive scheme which improves the rejection cost from…

Probability · Mathematics 2015-11-25 Richard Arratia , Stephen DeSalvo

Statistical inference is often simplified by sample-splitting. This simplification comes at the cost of the introduction of randomness not native to the data. We propose a simple procedure for sequentially aggregating statistics constructed…

Econometrics · Economics 2024-11-18 David M. Ritzwoller , Joseph P. Romano

We provide practical, efficient, and nonparametric methods for auditing the fairness of deployed classification and regression models. Whereas previous work relies on a fixed-sample size, our methods are sequential and allow for the…

Machine Learning · Statistics 2025-05-19 Ben Chugg , Santiago Cortes-Gomez , Bryan Wilder , Aaditya Ramdas

In this work we study the problem of inferring a discrete probability distribution using both expert knowledge and empirical data. This is an important issue for many applications where the scarcity of data prevents a purely empirical…

Machine Learning · Computer Science 2020-01-08 Rémi Besson , Erwan Le Pennec , Stéphanie Allassonnière

We consider the problem of accurate computation of the finite difference $f(\x+\s)-f(\x)$ when $\Vert\s\Vert$ is very small. Direct evaluation of this difference in floating point arithmetic succumbs to cancellation error and yields 0 when…

Optimization and Control · Mathematics 2013-07-17 Stephen Vavasis

Smoothing of noisy sample covariances is an important component in functional data analysis. We propose a novel covariance smoothing method based on penalized splines and associated software. The proposed method is a bivariate spline…

Methodology · Statistics 2017-04-07 Luo Xiao , Cai Li , William Checkley , Ciprian M. Crainiceanu

The r largest order statistics approach is widely used in extreme value analysis because it may use more information from the data than just the block maxima. In practice, the choice of r is critical. If r is too large, bias can occur; if…

Methodology · Statistics 2018-06-13 Brian Bader , Jun Yan , Xuebin Zhang

New methods for time-to-event prediction are proposed by extending the Cox proportional hazards model with neural networks. Building on methodology from nested case-control studies, we propose a loss function that scales well to large data…

Machine Learning · Statistics 2019-09-16 Håvard Kvamme , Ørnulf Borgan , Ida Scheel

With the growing availability of large-scale biomedical data, it is often time-consuming or infeasible to directly perform traditional statistical analysis with relatively limited computing resources at hand. We propose a fast subsampling…

Methodology · Statistics 2023-05-18 Haixiang Zhang , Lulu Zuo , HaiYing Wang , Liuquan Sun

Extreme multi-label classification (XML) is becoming increasingly relevant in the era of big data. Yet, there is no method for effectively generating stratified partitions of XML datasets. Instead, researchers typically rely on provided…

Machine Learning · Computer Science 2021-03-08 Maximillian Merrillees , Lan Du

Modern statistical analyses often encounter datasets with massive sizes and heavy-tailed distributions. For datasets with massive sizes, traditional estimation methods can hardly be used to estimate the extreme value index directly. To…

Methodology · Statistics 2022-07-26 Yongxin Li , Liujun Chen , Deyuan Li , Hansheng Wang

Nowadays, more and more datasets are stored in a distributed way for the sake of memory storage or data privacy. The generalized eigenvalue problem (GEP) plays a vital role in a large family of high-dimensional statistical models. However,…

Machine Learning · Computer Science 2021-11-25 Kexin Lv , Fan He , Xiaolin Huang , Jie Yang , Liming Chen

In this paper, a novel method for data splitting is presented: an iterative procedure divides the input dataset of volcanic eruption, chosen as the proposed use case, into two parts using a dissimilarity index calculated on the cumulative…

Computer Vision and Pattern Recognition · Computer Science 2024-10-10 Simona Reale , Pietro Di Stasio , Francesco Mauro , Alessandro Sebastianelli , Paolo Gamba , Silvia Liberata Ullo

Two-phase sampling offers a cost-effective way to validate error-prone covariate measurements in biomedical databases. Inexpensive or easy-to-obtain information is collected for the entire study in Phase I. Then, a subset of patients…

Methodology · Statistics 2026-05-21 Sarah C. Lotspeich , Cole Manschot

Private synthetic data sharing is preferred as it keeps the distribution and nuances of original data compared to summary statistics. The state-of-the-art methods adopt a select-measure-generate paradigm, but measuring large domain…

Cryptography and Security · Computer Science 2023-10-11 Meifan Zhang , Dihang Deng , Lihua Yin

Partition-wise models offer a flexible approach for modeling complex and multidimensional data that are capable of producing interpretable results. They are based on partitioning the observed data into regions, each of which is modeled with…

Methodology · Statistics 2017-06-07 Rex C. Y. Cheung , Alexander Aue , Thomas C. M. Lee

During the last few decades, online controlled experiments (also known as A/B tests) have been adopted as a golden standard for measuring business improvements in industry. In our company, there are more than a billion users participating…

Applications · Statistics 2021-08-06 Tao Xiong , Yihan Bao , Penglei Zhao , Yong Wang

Electronic voting systems have significant advantages in comparison with physical voting systems. One of the main challenges in e-voting systems is to secure the voting process: namely, to certify that the computed results are consistent…

Cryptography and Security · Computer Science 2025-05-21 Tamir Tassa , Lihi Dery , Arthur Zamarin

The most popular multiple testing procedures are stepwise procedures based on $P$-values for individual test statistics. Included among these are the false discovery rate (FDR) controlling procedures of Benjamini--Hochberg [J. Roy. Statist.…

Statistics Theory · Mathematics 2009-06-18 Arthur Cohen , Harold B. Sackrowitz , Minya Xu

This note discusses the problem of choosing between hypotheses in a situation with many, correlated non-normal variables. A new method is introduced to shrink the many variables into a smaller subset of variables with zero mean, unit…

Data Analysis, Statistics and Probability · Physics 2007-05-23 Byron P. Roe