English
Related papers

Related papers: Why Empirical p-Values Are Not Uniform: Reference …

200 papers

Quantifying uncertainty in detected changepoints is an important problem. However it is challenging as the naive approach would use the data twice, first to detect the changes, and then to test them. This will bias the test, and can lead to…

Methodology · Statistics 2026-05-11 Rachel Carrington , Paul Fearnhead

In a regression model, prediction is typically performed after model selection. The large variability in the model selection makes the prediction unstable. Thus, it is essential to reduce the variability in model selection and improve…

Computation · Statistics 2024-04-11 Wataru Yoshida , Kei Hirose

The estimation of individual treatment effects (ITE) focuses on predicting the outcome changes that result from a change in treatment. A fundamental challenge in observational data is that while we need to infer outcome differences under…

Machine Learning · Computer Science 2025-12-23 Zichuan Lin , Xiaokai Huang , Jiate Liu , Yuxuan Han , Jia Chen , Xiapeng Wu , Deheng Ye

A popular approach for testing if two univariate random variables are statistically independent consists of partitioning the sample space into bins, and evaluating a test statistic on the binned data. The partition size matters, and the…

Methodology · Statistics 2016-04-28 Ruth Heller , Yair Heller , Shachar Kaufman , Barak Brill , Malka Gorfine

The comparison of benchmark error sets is an essential tool for the evaluation of theories in computational chemistry. The standard ranking of methods by their Mean Unsigned Error is unsatisfactory for several reasons linked to the…

Methodology · Statistics 2020-09-29 Pascal Pernot , Andreas Savin

A sample covariance matrix $\boldsymbol{S}$ of completely observed data is the key statistic in a large variety of multivariate statistical procedures, such as structured covariance/precision matrix estimation, principal component analysis,…

Methodology · Statistics 2021-04-20 Seongoh Park , Xinlei Wang , Johan Lim

Compositional data (i.e., data comprising random variables that sum up to a constant) arises in many applications including microbiome studies, chemical ecology, political science, and experimental designs. Yet when compositional data serve…

Methodology · Statistics 2025-01-03 Ritwik Bhaduri , Siyuan Ma , Lucas Janson

The internal validity of observational study is often subject to debate. In this study, we define the counterfactuals as the unobserved sample and intend to quantify its relationship with the null hypothesis statistical testing (NHST). We…

Applications · Statistics 2019-06-21 Tenglong Li , Kenneth A. Frank

We analyze different types of simulations that applied researchers can use to assess whether their inference methods reliably control false-positive rates. We show that different assessments involve trade-offs, varying in the types of…

Econometrics · Economics 2025-10-03 Bruno Ferman

While the traditional viewpoint in machine learning and statistics assumes training and testing samples come from the same population, practice belies this fiction. One strategy -- coming from robust statistics and optimization -- is thus…

Machine Learning · Statistics 2024-07-08 Maxime Cauchois , Suyash Gupta , Alnur Ali , John C. Duchi

Statistical hypothesis testing and effect size measurement are routine parts of quantitative research. Advancements in computer processing power have greatly improved the capability of statistical inference through the availability of…

Methodology · Statistics 2024-01-18 Michael J. Crosse , John J. Foxe , Sophie Molholm

The paper proposes one-to-one transformation of the vector of components $\{Y_{in}\}_{i=1}^m$ of Pearson's chi-square statistic, \[Y_{in}=\frac{\nu_{in}-np_i}{\sqrt{np_i}},\qquad i=1,\ldots,m,\] into another vector $\{Z_{in}\}_{i=1}^m$,…

Statistics Theory · Mathematics 2014-01-06 Estate Khmaladze

Vanilla variational inference finds an optimal approximation to the Bayesian posterior distribution, but even the exact Bayesian posterior is often not meaningful under model misspecification. We propose predictive variational inference…

Machine Learning · Statistics 2026-03-31 Jinlin Lai , Antonio Linero , Yuling Yao

Recently, Saeb et al (2017) showed that, in diagnostic machine learning applications, having data of each subject randomly assigned to both training and test sets (record-wise data split) can lead to massive underestimation of the…

Split conformal prediction provides finite-sample marginal coverage under exchangeability, but this guarantee averages over the random calibration sample. We study instead the law of the calibration-conditional coverage induced by a…

Machine Learning · Statistics 2026-05-20 Thiago R. Ramos , Helton Graziadei , Luben M. C. Cabezas

Mutually uncorrelated random discrete events, manifesting a common basic process, are examined often in terms of their occurrence rate as a function of one or more of their distinguishing attributes, such as measurements of photon spectrum…

Instrumentation and Methods for Astrophysics · Physics 2012-07-26 Avinash A. Deshpande , Harsha Raichur

Pursuing invariant prediction from heterogeneous environments opens the door to learning causality in a purely data-driven way and has several applications in causal discovery and robust transfer learning. However, existing methods such as…

Statistics Theory · Mathematics 2025-01-30 Yihong Gu , Cong Fang , Yang Xu , Zijian Guo , Jianqing Fan

Safety-critical prediction systems, such as autonomous vehicles, weather forecasters, and medical monitors, commonly rely on probabilistic forecasters. These forecasters make predictions about possible future outcomes, and their quality and…

Methodology · Statistics 2026-04-30 Romeo Valentin

Distribution-free prediction sets play a pivotal role in uncertainty quantification for complex statistical models. Their validity hinges on reliable calibration data, which may not be readily available as real-world environments often…

Methodology · Statistics 2024-06-11 Elise Han , Chengpiao Huang , Kaizheng Wang

Binary endpoints are common in clinical trials and conditional odds ratios have traditionally been used to assess treatment effects. However, the interpretation of odds ratios is difficult, they are non-collapsible and rely on strong…

Methodology · Statistics 2026-05-20 Martin Schnuerch , Alex Ocampo , Klaus Kähler Holst , Christian Stock