Related papers: On some discrete random variables arising from rec…
This paper outlines a unified framework for high dimensional variable selection for classification problems. Traditional approaches to finding interesting variables mostly utilize only partial information through moments (like mean…
Variable selection comprises an important step in many modern statistical inference procedures. In the regression setting, when estimators cannot shrink irrelevant signals to zero, covariates without relationships to the response often…
This paper establishes complete convergence for weighted sums and the Marcinkiewicz--Zygmund-type strong law of large numbers for sequences of negatively associated and identically distributed random variables $\{X,X_n,n\ge1\}$ with general…
Neural networks and other machine learning models compute continuous representations, while humans communicate mostly through discrete symbols. Reconciling these two forms of communication is desirable for generating human-readable…
In this paper we study some asymptotic properties of the kernel conditional quantile estimator with randomly left-truncated data which exhibit some kind of dependence. We extend the result obtained by Lemdani, Ould-Sa\"id and Poulin [16] in…
As a popular tool for producing meaningful and interpretable models, large-scale sparse learning works efficiently when the underlying structures are indeed or close to sparse. However, naively applying the existing regularization methods…
For data living in a manifold $M\subseteq \mathbb{R}^m$ and a point $p\in M$ we consider a statistic $U_{k,n}$ which estimates the variance of the angle between pairs of vectors $X_i-p$ and $X_j-p$, for data points $X_i$, $X_j$, near $p$,…
Let $M_n$ be an $n\times n$ signed random combinatorial matrix whose rows are independent and uniformly distributed over the set of $\{-1,0,1\}$-vectors with exactly $n/2$ zero coordinates. Despite the dependence induced by the row…
Let $X=\{X_j , j\ge 1\}$ be a sequence of independent, square integrable variables taking values in a common lattice $\mathcal L(v_{ 0},D )= \{v_{ k}=v_{ 0}+D k , k\in \Z\}$. Let $S_n=X_1+\ldots +X_n$, $a_n= {\mathbb E\,} S_n$, and…
The problem is sequence prediction in the following setting. A sequence $x_1,...,x_n,...$ of discrete-valued observations is generated according to some unknown probabilistic law (measure) $\mu$. After observing each outcome, it is required…
This paper considers linear model selection when the response is vector-valued and the predictors are randomly observed. We propose a new approach that decouples statistical inference from the selection step in a "post-inference model…
Let $\{\xi_1,\xi_2,\ldots\}$ be a sequence of independent random variables, and $\eta$ be a counting random variable independent of this sequence. In addition, let $S_0:=0$ and $S_n:=\xi_1+\xi_2+\cdots+\xi_n$ for $n\geqslant1$. We consider…
Let $\xi_1, \xi_2,\ldots$ be a sequence of independent and identically distributed random variables with zero mean, finite second moment and regularly varying right distribution tail. Motivated by a stop-loss insurance model, we consider a…
Consider a sequence {X(i,0) : i = 1, ..., n} of i.i.d. random variables. Associate to each X(i,0) an independent mean-one Poisson clock. Every time a clock rings replace that X-variable by an independent copy. In this way, we obtain i.i.d.…
We present an efficient algorithm that, given a discrete random variable $X$ and a number $m$, computes a random variable whose support is of size at most $m$ and whose Kolmogorov distance from $X$ is minimal, also for the one-sided…
In unsupervised causal representation learning for sequential data with time-delayed latent causal influences, strong identifiability results for the disentanglement of causally-related latent variables have been established in stationary…
The general relationship between an arbitrary frequency distribution and the expectation value of the frequency distributions of its samples is discussed. A wide set of measurable quantities ("invariant moments") whose expectation value…
Consider $n$ players whose "scores" are independent and identically distributed values $\{X_i\}_{i=1}^n$ from some discrete distribution $F$. We pay special attention to the cases where (i) $F$ is geometric with parameter $p\to0$ and (ii)…
Balancing weights have been widely applied to single or monotone missingness due to empirical advantages over likelihood-based methods and inverse probability weighting approaches. This paper considers non-monotone missing data under the…
In Bayesian inference, we seek to compute information about random variables such as moments or quantiles on the basis of {available data} and prior information. When the distribution of random variables is {intractable}, Monte Carlo (MC)…