Related papers: A proposal for a different chi-square function for…
An important feature of Bayesian statistics is the opportunity to do sequential inference: the posterior distribution obtained after seeing a dataset can be used as prior for a second inference. However, when Monte Carlo sampling methods…
Regression for count data is widely performed by models such as Poisson, negative binomial (NB) and zero-inflated regression. A challenge often faced by practitioners is the selection of the right model to take into account dispersion,…
Context: Two-point correlation functions are used throughout cosmology as a measure for the statistics of random fields. When used in Bayesian parameter estimation, their likelihood function is usually replaced by a Gaussian approximation.…
If a discrete probability distribution in a model being tested for goodness-of-fit is not close to uniform, then forming the Pearson chi-square statistic can involve division by nearly zero. This often leads to serious trouble in practice…
As an application of Stein's method for Poisson approximation, we prove rates of convergence for the tail probabilities of two scan statistics that have been suggested for detecting local signals in sequences of independent random variables…
For integer valued random variables, the translated Poisson distributions form a flexible family for approximation in total variation, in much the same way that the normal family is used for approximation in Kolmogorov distance. Using the…
Mutual space-frequency distribution is proposed and it is shown that Wigner and Weyl distribution functions are only particular cases of these distribution. Mutual distribution for Gaussian signal is analytically obtained. The simple…
We develop a functional Stein-Malliavin method in a non-diffusive Poissonian setting, thus obtaining a) quantitative central limit theorems for approximation of arbitrary non-degenerate Gaussian random elements taking values in a separable…
Pearson's Chi-square test is a widely used tool for analyzing categorical data, yet its statistical power has remained theoretically underexplored. Due to the difficulties in obtaining its power function in the usual manner, Cochran (1952)…
New bounds on the total variation distance between the law of integer valued functionals of possibly non-symmetric and non-homogeneous infinite Rademacher sequences and the Poisson distribution are established. They are based on a…
This paper presents the application of a new semi-analytical method of linear regression for Poisson count data to COVID-19 events. The regression is based on the Bonamente and Spence (2022) maximum-likelihood solution for the best-fit…
The reduced chi-squared statistic is a commonly used goodness-of-fit measure, but it cannot easily detect features near the noise level, even when a large amount of data is available. In this paper, we introduce a new goodness-of-fit…
The C statistics, also known as the Cash statistic, is often used in astronomy for the analysis of low-count Poisson data. One of the challenges of the C statistic is that its probability distribution, under the null hypothesis that the…
Assume that we observe a large number of curves, all of them with identical, although unknown, shape, but with a different random shift. The objective is to estimate the individual time shifts and their distribution. Such an objective…
We address the issue of performing testing inference in generalized linear models when the sample size is small. This class of models provides a straightforward way of modeling normal and non-normal data and has been widely used in several…
This paper introduces a general method to approximate the convolution of an arbitrary program with a Gaussian kernel. This process has the effect of smoothing out a program. Our compiler framework models intermediate values in the program…
People employ the function-on-function regression to model the relationship between two random curves. Fitting this model, widely used strategies include algorithms falling into the framework of functional partial least squares (typically…
As a powerful tool for longitudinal data analysis, the generalized estimating equations have been widely studied in the academic community. However, in large-scale settings, this approach faces pronounced computational and storage…
Time-dependent ensemble averages, i.e., trajectory-based averages of some observable, are of importance in many fields of science. A crucial objective when interpreting such data is to fit these averages (for instance, squared…
Statistical divergences are ubiquitous in machine learning as tools for measuring discrepancy between probability distributions. As these applications inherently rely on approximating distributions from samples, we consider empirical…