Related papers: Probability that a chromosome is lost without trac…
An analysis of the influence of missing samples in signals exhibiting sparsity in the Hermite transform domain is provided. Based on the statistical properties derived for the Hermite coefficients of randomly undersampled signal, the…
Amidst rising appreciation for privacy and data usage rights, researchers have increasingly acknowledged the principle of data minimization, which holds that the accessibility, collection, and retention of subjects' data should be kept to…
A discrete-time totally asymmetric simple exclusion process on a lattice with open boundaries is considered. There are particles of different types. The type of a particle is characterized by the probability that a particle moves to a…
We develop a framework for approximating collapsed Gibbs sampling in generative latent variable cluster models. Collapsed Gibbs is a popular MCMC method, which integrates out variables in the posterior to improve mixing. Unfortunately for…
Trace reconstruction is the problem of learning an unknown string $x$ from independent traces of $x$, where traces are generated by independently deleting each bit of $x$ with some deletion probability $q$. In this paper, we initiate the…
This research deals with the estimation and imputation of missing data in longitudinal models with a Poisson response variable inflated with zeros. A methodology is proposed that is based on the use of maximum likelihood, assuming that data…
We consider a lattice-inspired random matrix model for the QCD chiral phase transition at finite chemical potential. Useful features of the usual RMM for QCD at finite chemical potential are reobtained, some being brought closer to their…
Let $W_1,\ldots,W_N$ be a sample of $\mathrm{Pareto}(\alpha)$ random variables normalized by their sum, such that $\sum_i W_i=1$. The $W_i$ may represent the weights of valleys in a spin glass (if $0<\alpha<1$), or the frequency of…
In the trace reconstruction problem our goal is to learn an unknown string $x\in \{0,1\}^n$ given independent traces of $x$. A trace is obtained by independently deleting each bit of $x$ with some probability $\delta$ and concatenating the…
Matrix permanent plays a key role in data association probability calculations. Exact algorithms (such as Ryser's) scale exponentially with matrix size. Fully polynomial time randomized approximation schemes exist but are quite complex.…
Within the universal zero-range theory, we compute the three-body recombination rate to deep molecular states for two identical bosons resonantly interacting with each other and with a third atom of another species, in the absence of weakly…
In recent years, several algorithms, which approximate matrix decomposition, have been developed. These algorithms are based on metric conservation features for linear spaces of random projection types. We show that an i.i.d sub-Gaussian…
We consider a population with two types of individuals, distinguished by the resources required for reproduction: type-$0$ (small) individuals need a fractional resource unit of size $\vartheta \in (0,1)$, while type-$1$ (large) individuals…
There are numerous examples of natural and artificial processes that represent stochastic sequences of events followed by an absolute refractory period during which the occurrence of a subsequent event is impossible. In the simplest case of…
In many real-world applications of machine learning classifiers, it is essential to predict the probability of an example belonging to a particular class. This paper proposes a simple technique for predicting probabilities based on…
In the trace reconstruction problem, one attempts to reconstruct a fixed but unknown string $x$ of length $n$ from a given number of traces $\tilde{x}$ drawn iid from the application of a noisy process (such as the deletion channel) to $x$.…
We present a calculation technique for modeling inhomogeneous DNA replication kinetics, where replication factors such as initiation rates or fork speeds can change with both position and time. We can use our model to simulate data sets…
Randomized experiments are a crucial tool for causal inference in many different fields. Rerandomization addresses any covariate imbalance in such experiments by resampling treatment assignments until certain balance criteria are satisfied.…
Fitting high-dimensional statistical models often requires the use of non-linear parameter estimation procedures. As a consequence, it is generally impossible to obtain an exact characterization of the probability distribution of the…
We present herein a scheme by which to accurately evaluate the error exponents of a lossy data compression problem, which characterize average probabilities over a code ensemble of compression failure and success above or below a critical…