Statistics
Online platforms prioritize long-term business outcomes, yet typical experiments are far too short to measure these outcomes directly. Our goal in this paper is to collect and share industry knowledge on how to make decisions from…
As machine learned models increase in complexity and expressive power, features of simpler models, such as interpretability and control over the shape of the modeled function are lost. On the one edge of the spectrum we have simple linear…
Testing independence in an R by C table is usually posed as one omnibus statistic against a dense alternative. When the departure is sparse, concentrated in a few cells, the sharp object is instead a detection boundary: the weakest signal…
Simulating container transshipment hubs requires vessel arrivals reflecting cyclical liner schedules and origin-destination (OD) cargo pairing. Existing models relying on Poisson arrivals and aggregate transshipment volumes severely distort…
Accurate, up-to-date income data at the sub-municipal scale is essential for social policy in middle-income countries, yet in Brazil it depends on a costly decennial census whose intercensal gap recently exceeded a decade. We test whether…
The Cox model remains the default for survival analysis, but the proportional hazards assumption is often violated and hazard ratios can be difficult to interpret. Accelerated failure time (AFT) models provide an intuitive time-scale…
Estimating individualized treatment regimes (ITRs) is fundamental in data-driven personalized decision-making problems, such as precision medicine. Most of the ITR literature either focuses on categorical/continuous treatments or assumes no…
Prediction intervals for multi-modal regression with tabular variables, text, images, or other input sources are difficult to calibrate when those sources disagree or one is missing. A single global quantile averages these regimes together…
Large-scale assessments suppress subgroup achievement estimates below minimum sample size thresholds, such as the National Assessment of Educational Progress (NAEP) rule of 62, disproportionately affecting historically underrepresented…
These days, Sri Lanka is facing a severe dengue outbreak. In this situation, it is very useful to provide an overall understanding of Sri Lanka's historical dengue trends as well as the current dengue situation. A dashboard is an easy and…
Sampling high-dimensional probability distributions is a central task in scientific computing, with applications ranging from Bayesian inference to statistical physics and molecular simulation. Despite decades of methodological…
Surrogate endpoints are widely used in clinical trials to accelerate treatment evaluation, yet their validity may vary substantially across patient subgroups. Although recent advances in heterogeneous causal mediation analysis enable…
Stochastic gradient descent for a loss function discontinuous across lower dimensional manifolds is analyzed by studying its differential equation limit.
This study investigated the behavior of Principal Component Analysis (PCA) when applied to datasets with extremely large numbers of observations. Although statistical theory suggests that sampling error diminishes and sample estimates…
A growing literature argues that progress in human longevity is slowing, and that this slowdown may signal an approaching biological limit to lifespan. The argument rests on visible flattening of cumulative life expectancy records,…
Hierarchical composite endpoints are commonly used in clinical trials when component out- comes differ in clinical importance. The win ratio compares patients across treatment groups according to a pre-specified order of clinical priority…
Understanding the relationship between home and away goal counts in football provides valuable insights into match-level dynamics. While the influence of home advantage is well-established, with historical records indicating roughly 50% of…
We consider the optimization of the Optimized Certainty Equivalent (OCE) risk, with applications including portfolio optimization in finance, and uncertainty quantification, classification, and regression in machine learning. Our…
This paper introduces Mixtures of Geodesic Factor Analyzers (MGFA) on Riemannian homogeneous spaces. MGFA uses a geodesic factor model within each mixture component, providing greater expressiveness than mixtures of Riemannian radial…
Item response models usually assume a normal trait distribution, yet little is known about how often fitted distributions in real studies differ substantially from normality or which reported results are most affected. We analyzed 504…