English
Related papers

Related papers: Accurate $p$-Value Calculation for Generalized Fis…

200 papers

Gaussian processes (GPs) are generally regarded as the gold standard surrogate model for emulating computationally expensive computer-based simulators. However, the problem of training GPs as accurately as possible with a minimum number of…

Methodology · Statistics 2024-11-26 Hossein Mohammadi , Peter Challenor

Modern social and biomedical scientific publications require the reporting of covariate balance tables with not only covariate means by treatment group but also the associated $p$-values from significance tests of their differences. The…

Methodology · Statistics 2023-06-06 Anqi Zhao , Peng Ding

A novel method for computing exact p-values of one-sided statistics from the Kolmogorov-Smirnov family is presented. It covers the Higher Criticism statistic, one-sided weighted Kolmogorov-Smirnov statistics, and the one-sided Berk-Jones…

Computation · Statistics 2023-08-14 Amit Moscovich

Multi-sensor data fusion technology plays an important role in real applications. Because of the flexibility and effectiveness in modelling and processing the uncertain information regardless of prior probabilities, Dempster-Shafer evidence…

Artificial Intelligence · Computer Science 2018-06-06 Fuyuan Xiao

Bayesian predictive probabilities are commonly used for interim monitoring of clinical trials through efficacy and futility stopping rules. Despite their usefulness, calculation of predictive probabilities, particularly in pre-experiment…

Applications · Statistics 2024-06-18 Joe Marion , Liz Lorenzi , Cora Allen-Savietta , Scott Berry , Kert Viele

The test of independence is a crucial component of modern data analysis. However, traditional methods often struggle with the complex dependency structures found in high-dimensional data. To overcome this challenge, we introduce a novel…

Methodology · Statistics 2024-09-13 Mingshuo Liu , Doudou Zhou , Hao Chen

Adaptive experiments use preliminary analyses of the data to inform further course of action and are commonly used in many disciplines including medical and social sciences. Because the null hypothesis and experimental design are…

Methodology · Statistics 2026-05-26 Tobias Freidling , Qingyuan Zhao , Zijun Gao

Testing cross-sectional independence in panel data models is of fundamental importance in econometric analysis with high-dimensional panels. Recently, econometricians began to turn their attention to the problem in the presence of serial…

Methodology · Statistics 2023-09-18 Hongfei Wang , Binghui Liu , Long Feng , Yanyuan Ma

Quantifying the influence of infinitesimal changes in training data on model performance is crucial for understanding and improving machine learning models. In this work, we reformulate this problem as a weighted empirical risk minimization…

Machine Learning · Computer Science 2025-04-11 Omri Lev , Ashia C. Wilson

Motivation: Combining the results of different experiments to exhibit complex patterns or to improve statistical power is a typical aim of data integration. The starting point of the statistical analysis often comes as sets of p-values…

Methodology · Statistics 2021-12-02 Tristan Mary-Huard , Sarmistha Das , Indranil Mukhopadhyay , Stéphane Robin

In a modern observational study based on healthcare databases, the number of observations and of predictors typically range in the order of $10^5$ ~ $10^6$ and of $10^4$ ~ $10^5$. Despite the large sample size, data rarely provide…

Computation · Statistics 2022-03-30 Akihiko Nishimura , Marc A. Suchard

In a recent simulation study, Goodman et al. (2019) compare several methods with regard to their type I and type II error rates in case of a thick null hypothesis that includes all values that are practically equivalent to the point null…

Methodology · Statistics 2022-06-07 Robin Tim Dreher , Leona Hoffmann , Arne Kramer-Sunderbrink , Peter Pütz , Robin Werner

Many statistical models require an estimation of unknown (co)-variance parameter(s) in a model. The estimation usually obtained by maximizing a log-likelihood which involves log determinant terms. In principle, one requires the…

Computation · Statistics 2016-09-05 Shengxin Zhu , Tongxiang Gu , Xiaowen Xu , Zeyao Mo

Predictions and generations from large language models are increasingly being explored as an aid in limited data regimes, such as in computational social science and human subjects research. While prior technical work has mainly explored…

Machine Learning · Computer Science 2025-10-09 Yewon Byun , Shantanu Gupta , Zachary C. Lipton , Rachel Leah Childers , Bryan Wilder

This work proposes a new method for computing acceptance regions of exact multinomial tests. From this an algorithm is derived, which finds exact p-values for tests of simple multinomial hypotheses. Using concepts from discrete convex…

Computation · Statistics 2023-05-31 Johannes Resin

Discrete state spaces represent a major computational challenge to statistical inference, since the computation of normalisation constants requires summation over large or possibly infinite sets, which can be impractical. This paper…

Methodology · Statistics 2023-09-04 Takuo Matsubara , Jeremias Knoblauch , François-Xavier Briol , Chris. J. Oates

In an attempt to provide an answer to the increasing criticism against p-values and to bridge the gap between statistical inference and prediction modelling, we introduce the probability of improved prediction (PIP). In general, the PIP is…

Methodology · Statistics 2024-05-28 Olivier Thas , Stijn Jaspers

Many testing problems are readily amenable to randomised tests such as those employing data splitting. However despite their usefulness in principle, randomised tests have obvious drawbacks. Firstly, two analyses of the same dataset may…

Methodology · Statistics 2024-09-05 F. Richard Guo , Rajen D. Shah

In randomized experiments, treatment and control groups should be roughly the same--balanced--in their distributions of pretreatment variables. But how nearly so? Can descriptive comparisons meaningfully be paired with significance tests?…

Methodology · Statistics 2008-08-29 Ben B. Hansen , Jake Bowers

We study the D-optimal Data Fusion (DDF) problem, which aims to select new data points, given an existing Fisher information matrix, so as to maximize the logarithm of the determinant of the overall Fisher information matrix. We show that…

Optimization and Control · Mathematics 2022-08-09 Yongchun Li , Marcia Fampa , Jon Lee , Feng Qiu , Weijun Xie , Rui Yao