Related papers: The Poisson Multinomial Distribution and Its Appli…
In this paper, we propose a discrete circular distribution obtained by extending the wrapped Poisson distribution. This new distribution, the Invariant Wrapped Poisson (IWP), enjoys numerous advantages: simple tractable density,…
Sample- and computationally-efficient distribution estimation is a fundamental tenet in statistics and machine learning. We present SURF, an algorithm for approximating distributions by piecewise polynomials. SURF is: simple, replacing…
Distribution matching and dematching (DM/invDM) are key functions in probabilistic shaping (PS). Recently techniques for low complexity implementation of DM/invDM have been well studied. Our previously proposed hierarchical DM (HiDM) is one…
Multi-parton distributions in a proton, the nonperturbative quantities needed to make predictions for multiple scattering rates, are poorly constrained from theory and data and must be modelled. All Monte Carlo event generators that…
Covariance matrix estimation is an important problem in multivariate data analysis, both from theoretical as well as applied points of view. Many simple and popular covariance matrix estimators are known to be severely affected by model…
Probability distributions produced by the cross-entropy loss for ordinal classification problems can possess undesired properties. We propose a straightforward technique to constrain discrete ordinal probability distributions to be unimodal…
Be d_{m,n} a generic element in the infinite matrix D, with d_{1, n} defined as the n-th prime number and, for any m>1, d_{m, n} = | d_{m-1, n} - d_{m-1, n+1} | When n>1, after the first few terms the columns in the matrix appear to be…
A well-known challenge in the semantics of programming languages is how to combine non-determinism and probability. At a technical level, the problem arises from the fact that there is a no distributive law between the powerset monad and…
A key element in transfer learning is representation learning; if representations can be developed that expose the relevant factors underlying the data, then new tasks and domains can be learned readily based on mappings of these salient…
In this letter, we derive the CDF (cumulative distribution function) of $k$th contact distance (CD) and nearest neighbor distance (NND) of the $n$-dimensional ($n$-D) Mat\'ern cluster process (MCP). We present a new approach based on the…
The paper investigates the problem of performing correlation analysis when the number of observations is very large. In such a case, it is often necessary to combine the random observations to achieve dimensionality reduction of the…
The statistical distribution of levels of an integrable system is claimed to be a Poisson distribution. In this paper, we numerically generate an ensemble of N dimensional random diagonal matrices as a model for regular systems. We evaluate…
Stochastic dominance (SD) provides a quantile-based partial ordering of random variables and has broad applications. Its extension to multivariate settings, however, is challenging due to the lack of a canonical ordering in $\mathbb{R}^d$…
Cumulative probability models (CPMs) are a robust alternative to linear models for continuous outcomes. However, they are not feasible for very large datasets due to elevated running time and memory usage, which depend on the sample size,…
We present an analysis of parton distribution functions (PDFs) of the proton using Markov Chain Monte Carlo (MCMC) methods. The MCMC approach naturally implements Bayes' theorem and thus provides a means to directly sample the underlying…
This paper studies fundamental aspects of modelling data using multivariate Watson distributions. Although these distributions are natural for modelling axially symmetric data (i.e., unit vectors where $\pm \x$ are equivalent), for…
Given a set of independent Poisson random variables with common mean, we study the distribution of their maximum and obtain an accurate asymptotic formula to locate the most probable value of the maximum. We verify our analytic results with…
Nonuniform subsampling methods are effective to reduce computational burden and maintain estimation efficiency for massive data. Existing methods mostly focus on subsampling with replacement due to its high computational efficiency. If the…
The compound decision problem for a vector of independent Poisson random variables with possibly different means has half a century old solution. However, it appears that the classical solution needs smoothing adjustment even when there are…
Planning and learning in Partially Observable MDPs (POMDPs) are among the most challenging tasks in both the AI and Operation Research communities. Although solutions to these problems are intractable in general, there might be special…