Related papers: Fully explicit large deviation inequalities for em…
A fundamental algorithm for selecting ranks from a finite subset of an ordered set is Radix Selection. This algorithm requires the data to be given as strings of symbols over an ordered alphabet, e.g., binary expansions of real numbers. Its…
We consider the residual empirical process in random design regression with long memory errors. We establish its limiting behaviour, showing that its rates of convergence are different from the rates of convergence for to the empirical…
In this note we use an approximation scheme to establish large deviations for quasi-periodic Gevrey cocycles. As an application, we obtain continuity in the cocycle for the Lyapunov exponent.
We develop a qualitative theory of Markov Decision Processes (MDPs) and Partially Observable MDPs that can be used to model sequential decision making tasks when only qualitative information is available. Our approach is based upon an…
Vapnik-Chervonenkis (VC) dimension is a fundamental measure of the generalization capacity of learning algorithms. However, apart from a few special cases, it is hard or impossible to calculate analytically. Vapnik et al. [10] proposed a…
In inverse problems, one attempts to infer spatially variable functions from indirect measurements of a system. To practitioners of inverse problems, the concept of "information" is familiar when discussing key questions such as which parts…
The purpose of this study is to introduce a new approach to feature ranking for classification tasks, called in what follows greedy feature selection. In statistical learning, feature selection is usually realized by means of methods that…
In many applications of relational learning, the available data can be seen as a sample from a larger relational structure (e.g. we may be given a small fragment from some social network). In this paper we are particularly concerned with…
Let $F$ be a class of functions on a probability space $(\Omega,\mu)$ and let $X_1,...,X_k$ be independent random variables distributed according to $\mu$. We establish high probability tail estimates of the form $\sup_{f \in F} |\{i :…
While effective concentration inequalities for suprema of empirical processes exist under boundedness or strict tail assumptions, no comparable results have been available under considerably weaker assumptions. In this paper, we derive…
A survey is given of some Chernoff type bounds for the tail probabilities P(X-EX > a) and P(X-EX < a) when X is a random variable that can be written as a sum of indicator variables that are either independent or negatively related. Most…
We consider a family of positive operator valued measures associated with representations of compact connected Lie groups. For many independent copies of a single state and a tensor power representation we show that the observed probability…
Divergences often play important roles for study in information science so that it is indispensable to investigate their fundamental properties. There is also a mathematical significance of such results. In this paper, we introduce some…
We establish a Large Deviations Principle for stochastic processes with Lipschitz continuous oblique reflections on regular domains. The rate functional is given as the value function of a control problem and is proved to be good. The proof…
Asymptotics deviation probabilities of the sum S n = X 1 + $\times$ $\times$ $\times$ + X n of independent and identically distributed real-valued random variables have been extensively investigated, in particular when X 1 is not…
We investigate the small deviation probabilities of a class of very smooth stationary Gaussian processes playing an important role in Bayesian statistical inference. Our calculations are based on the appropriate modification of the entropy…
Multivariate (or vector-valued) processes are important for modeling multiple variables. The fractal indices of the components of the underlying multivariate process play a key role in characterizing the dependence structures and…
We study data processing inequalities that are derived from a certain class of generalized information measures, where a series of convex functions and multiplicative likelihood ratios are nested alternately. While these information…
We study a separable design for computing information measures, where the information measure is computed from learned feature representations instead of raw data. Under mild assumptions on the feature representations, we demonstrate that a…
When dealing with modern big data sets, a very common theme is reducing the set through a random process. These generally work by making "many simple estimates" of the full data set, and then judging them as a whole. Perhaps magically,…