Statistics
Nearly every quantitative study in psychology reports a reliability coefficient, so the field knows a great deal about the reliability that authors choose to publish. It knows much less about the reliability of the data psychology actually…
Covariate adjustment improves the efficiency of treatment-effect analyses in randomized clinical trials, provided the adjustment targets the correct quantity. For time-to-event endpoints, two marginal targets are of primary interest: the…
Polypharmacy, commonly defined as the concurrent use of multiple medications, is increasingly prevalent in aging populations and is associated with adverse health outcomes. Motivated by longitudinal studies of aging people with HIV (PWH),…
Medical imaging provides rich information for predicting Alzheimer's disease progression, but existing imaging-based methods typically focus on a single survival endpoint and treat transition times as exactly observed or right-censored.…
Latent space models represent network nodes as points in a geometric space, with connection probabilities determined by distances between latent positions under a fixed metric, typically Euclidean, spherical, or hyperbolic. We generalize…
Statistical inference on individual activity networks has been a historically difficult task due to the lack of available data at the appropriate granularity and the complexity of modeling individual mobility patterns. The recent…
High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional…
Studying consequences following baseline exposures has become increasingly important for advancing comparative effectiveness research using real-world data. This case study evaluates the impact of hospital-acquired conditions (HAC) during…
Athletic careers yield sparse, irregular longitudinal series: few seasons per athlete, incomplete paths, and selection into continued play. Scientific interest often centres on proximity to peak attainable performance-a ceiling-rather than…
Classical likelihood-ratio tests and $\Delta$AIC exacerbate the statistical significance crisis by scaling with sample size, often flagging negligible improvements as highly significant. While causal estimands like the average treatment…
Random-projection tests for functional data depend on the probability law used to generate projection directions. A measure has to be selected to generate the random directions in which the data is projected. In $L^2$, probability measures…
Multinational HIV cohort studies face regulatory barriers to cross-border sharing of individual participant data, limiting centralized pooled analyses. Federated statistical methods, which exchange only aggregated information, offer a…
Latent trajectory analysis is a statistical method for explaining heterogeneity by partitioning patients into homogeneous subgroups based on similarities in outcome variables. In the context of clinical work, patients often do not follow…
Latent subgroup analysis is central to fields such as genomics, precision medicine, and social science, where the goal is to identify heterogeneous populations with distinct covariate structures or response behaviors. Mixture models provide…
Vector autoregressive moving-average (VARMA) models have long been considered impractical beyond moderate dimensions: the likelihood is non-convex, the parametrization is identified only up to equivalence, and every evaluation costs a pass…
A monotone adversary observes an i.i.d. labeled sample and appends a finite number of further examples of its choice, every one of them labeled correctly by the target hypothesis. The learner sees a uniform shuffle of the combined sample…
The autocorrelation function (ACF) and partial autocorrelation function (PACF) are foundational tools for identifying autoregressive moving-average (ARMA) models, yet they are often introduced to students as computational recipes…
We propose a new unit of analysis for longitudinal data: the Latent Memory Table. The scientific contribution is not the encoder. It is that table, treated as a reusable statistical object on the same footing as a matrix of…
Persistence diagrams (PDs) provide stable and interpretable summaries of multiscale topological structure. While substantial progress has been made in the statistical analysis of PDs, existing literature often treats diagrams as static…
In overparameterised classification, training data can be linearly separable even when the underlying distribution is not. In this setting, gradient descent (GD) on the logistic loss diverges in norm while converging in direction to a…