Related papers: On Identifying a Massive Number of Distributions
We study a special case of the problem of statistical learning without the i.i.d. assumption. Specifically, we suppose a learning method is presented with a sequence of data points, and required to make a prediction (e.g., a classification)…
Denote by $\mathbb{N}$ and $\mathbb{P}$ the set of all positive integers and prime numbers, respectively. Let $\mathbb{P}=\{p_1<p_2<\dots <p_n<\dots\}$, where $p_n$ is the $n$-th prime number. For $k\in\mathbb{N}$ we recursively define…
Commonly observed patterns typically follow a few distinct families of probability distributions. Over one hundred years ago, Karl Pearson provided a systematic derivation and classification of the common continuous distributions. His…
Let X_1,...., X_n be a collection of iid discrete random variables, and Y_1,..., Y_m a set of noisy observations of such variables. Assume each observation Y_a to be a random function of some a random subset of the X_i's, and consider the…
The distribution of differences of consecutive members of sequences of primes is investigated. A quantitative measure for oscillations among these differences is the curvature of the sequence. If the sequence is not too sparse, then sharp…
Given only aggregate choice data and limited information about how menus are distributed across the population, we describe what can be inferred robustly about the distribution of preferences (or more general decision rules). We strengthen…
Given a set of independent Poisson random variables with common mean, we study the distribution of their maximum and obtain an accurate asymptotic formula to locate the most probable value of the maximum. We verify our analytic results with…
We consider the estimation of the mixing distribution of a normal distribution where both the shift and scale are unobserved random variables. We argue that in general, the model is not identifiable. We give an elegant non-constructive…
The problem of reconstructing a sequence of independent and identically distributed symbols from a set of equal size, consecutive, fragments, as well as a dependent reference sequence, is considered. First, in the regime in which the…
The relationship between three probability distributions and their maximizable entropy forms is discussed without postulating entropy property. For this purpose, the entropy I is defined as a measure of uncertainty of the probability…
We consider the best-choice problem for independent (not necessarily iid) observations $X_1, \cdots, X_n$ with the aim of selecting the sample minimum. We show that in this full generality the monotone case of optimal stopping holds and the…
The distribution of frequency counts of distinct words by length in a language's vocabulary will be analyzed using two methods. The first, will look at the empirical distributions of several languages and derive a distribution that…
Multiplicity distributions, P(N), provide valuable information on the mechanism of the production process. We argue that the observed P(N) contain more information (located in the small N region) than expected and used so far. We…
We study the best-choice problem for processes which generalise the process of records from Poisson-paced i.i.d. observations. Under the assumption that the observer knows distribution of the process and the horizon, we determine the…
To infer a diffusion network based on observations from historical diffusion processes, existing approaches assume that observation data contain exact occurrence time of each node infection, or at least the eventual infection statuses of…
Pyrosequencing is emerging as one of the important next-generation sequencing technologies. We derive the statistical distributions of this technique in terms of nucleotide probabilities of the target sequences. We give exact distributions…
We investigate the probability of observing a given pattern of $n$ rises and falls in a random stationary data series. The data are modelled as a sequence of $n+1$ independent and identically distributed random numbers. This probabilistic…
With reference to a previous work, the problem of the experimental detection of non-causal synordination patterns between two series of physical events is examined. It is necessary that the patterns in question act in a reproducible, or at…
We consider the problem of hypothesis testing for discrete distributions. In the standard model, where we have sample access to an underlying distribution $p$, extensive research has established optimal bounds for uniformity testing,…
In counting experiments, one can set an upper limit on the rate of a Poisson process based on a count of the number of events observed due to the process. In some experiments, one makes several counts of the number of events, using…