Statistics
Beta diversity quantifies variation in species composition across ecological communities and is fundamental for understanding biodiversity patterns across space and environmental gradients. Statistical inference on beta diversity is…
There exist tests calibrated under a null narrower than the one implied by their test statistic -- the detection-null set. The part of the detection-null set not in the null is the test's blind spot. A framework for assessing detection…
Conditional Independence (CI) tests are the statistical engine of constraint-based causal discovery: in algorithms such as PC (Peter-Clark) and FCI (Fast Causal Inference), skeleton pruning and key orientations follow directly from CI…
Assessing the goodness-of-fit of a logistic regression model is a critical prerequisite before the model is used for inference. However, goodness-of-fit (GOF) tests such as the chi-square and deviance tests often give invalid results when…
The simplicial depth (SD) is a commonly used indicator of the centrality of points $x\in\mathbb{R}^d$ with respect to distributions $P$ on $\mathbb{R}^d$. Asymptotic theory for the sample SD is based on its representation as a…
Generative artificial intelligence (GenAI) is a large language model (LLM) that has the ability to generate media based on user-provided prompts. Given the demonstrated capabilities of models such as ChatGPT in information synthesis and…
Research on the use of medications during pregnancy has two primary goals: to detect signals that medications may be harmful to a pregnant individual or fetus, and to support better treatment of pregnant people who require pharmacotherapy.…
Reliable estimation of cure fractions depends critically on adequate follow-up. Classical procedures for assessing follow-up sufficiency are primarily inferential, whereas a separate body of literature estimates time to cure using…
The covariate-adjusted log-rank test is a novel method for covariate adjustment in randomized trials with time-to-event endpoints, offering guaranteed efficiency gains compared to the standard log-rank test. However, it has been noted that,…
Multidomain assessment batteries generate ordinal item responses that are often summarized through latent attribute profiles. Conventional cognitive diagnostic models (CDMs) provide interpretable measurement models for such profiles, but…
Constant-stepsize temporal-difference (TD) learning is attractive for policy evaluation, but inference from a single Markov trajectory must account for serial dependence and a stepsize-dependent stationary target. For fixed-stepsize linear…
Spectral clustering methods for network data are commonly based on a few matrix representations, such as the adjacency matrix and the symmetric Laplacian. We study a continuum of degree-normalized spectral embeddings that includes these…
Optimal stratification aggregates strata into a small number of final strata to minimise total sample size required to meet target precision constraints. This combinatorial objective, reformulated as a within-cluster dispersion surrogate,…
Whether an additional general dimension is necessary beyond correlated first-order factors is a property of the population covariance matrix, not of any estimator or design. This research establishes when that property can be decided. Where…
Prioritized pairwise outcomes are useful when clinical events follow a natural hierarchy, but censoring before pair resolution complicates estimation. We develop a win-ratio regression framework for this setting by defining a complete-data…
Conventional mean-only regression models are often too restrictive for the analysis of complex survey data, where interest frequently extends beyond the conditional mean to other aspects of the response distribution. Generalised additive…
Uncertainty in wind farm layout optimization regarding model choice and the impact of global blockage always exists. This paper evaluates these interactions by extending a multi-objective approach that maximizes mean Annual Energy…
Many forecasting systems produce point forecasts even when decisions require information about uncertainty. We investigate whether post-processing methods can systematically improve upon traditional Gaussian predictive distributions…
Predictively oriented (PrO) inference quantifies uncertainty by selecting a distribution over model parameters to optimize a scoring rule applied to the induced predictive distribution, together with a divergence penalty from a reference…
How many directions does a neural representation use to encode a concept? A common answer repeatedly erases probe directions and reports the stopping count or cumulative removed rank. We show that both quantities can change under an…