相关论文: Sharp Threshold for the Convergence of Nonstationa…
Given $N\ge2$ closed subspaces $M_1,\dotsc, M_N$ of a Hilbert space $X$, let $P_k$ denote the orthogonal projection onto $M_k$, $1\le k\le N$. It is known that the sequence $(x_n)$, defined recursively by $x_0=x$ and $x_{n+1}=P_N\cdots…
Neural networks have been very successful in many applications; we often, however, lack a theoretical understanding of what the neural networks are actually learning. This problem emerges when trying to generalise to new data sets. The…
Although overparameterized models have achieved remarkable practical success, their theoretical properties, particularly their generalization behavior, remain incompletely understood. The well known double descents phenomenon suggests that…
Let $p_n(y)=\sum_k\hat{\alpha}_k\phi(y-k)+\sum_{l=0}^{j_n-1}\sum_k\hat {\beta}_{lk}2^{l/2}\psi(2^ly-k)$ be the linear wavelet density estimator, where $\phi$, $\psi$ are a father and a mother wavelet (with compact support),…
We consider kernel smoothed Grenander-type estimators for a monotone hazard rate and a monotone density in the presence of randomly right censored data. We show that they converge at rate $n^{2/5}$ and that the limit distribution at a fixed…
We characterize sequences of numbers $(a_n)$ such that $\sum_{n\geq 1} a_n\Phi_n$ converges a.e. for any orthonormal system $(\Phi_n)$ in any $L_2$-space. In our criterion, we use the set $B =\{\sum_{m\geq n} |a_m|^2; n\geq 1\}$ and its…
The effect of errors in variables in empirical minimization is investigated. Given a loss $l$ and a set of decision rules $\mathcal{G}$, we prove a general upper bound for an empirical minimization based on a deconvolution kernel and a…
The threshold network model is a type of finite random graphs. In this paper, we introduce a generalized threshold network model. A pair of vertices with random weights is connected by an edge when real-valued functions of the pair of…
Finite mixtures are a cornerstone of Bayesian modelling, and it is well-known that sampling from the resulting posterior distribution can be a hard task. In particular, popular reversible Markov chain Monte Carlo schemes are often slow to…
In this paper we propose a neural network model with a novel Sequential Attention layer that extends soft attention by assigning weights to words in an input sequence in a way that takes into account not just how well that word matches a…
This paper presents a novel approach to address the constrained coding challenge of generating almost-balanced sequences. While strictly balanced sequences have been well studied in the past, the problem of designing efficient algorithms…
We consider the Lame system of linear elasticity with periodically distributed inclusions whose elastic parameters have high contrast compared to the background media. We develop a unified method based on layer potential techniques to…
In the averaging process on a graph $G = (V, E)$, a random mass distribution $\eta$ on $V$ is repeatedly updated via transformations of the form $\eta_{v}, \eta_{w} \mapsto (\eta_{v} + \eta_{w})/2$, with updates made according to…
Action convergence provides a limit theory for linear bounded operators $A_n:L^{\infty}(\Omega_n)\longrightarrow L^1(\Omega_n)$ where $\Omega_n$ are potentially different probability spaces. This notion of convergence emerged in graph…
Let $\{X, X_n, n\geq 1\}$ be a sequence of independent identically distributed non-degenerate random variables. Put $S_0=0, S_n = \sum^n_{i=1} X_i$ and $V_n^2=\sum^n_{i=1} X_i^2, n\ge 1.$ A weak convergence theorem is established for the…
We study the fine-grained uniform convergence behavior of halfspaces beyond worst-case VC bounds. For inhomogeneous halfspaces in $\mathbb{R}^d$ with $d\ge 2$, we show that standard first-order VC bounds are essentially tight: even…
In this note we establish a uniform bound for the distribution of a sum $S_n=X_1+\cdots+X_n$ of independent non-homogeneous Bernoulli trials. Specifically, we prove that $\sigma_n \mathbb{P}(S_n\!=\!j)\leq\eta$ where $\sigma_n$ denotes the…
For positive integers $r > \ell$, an $r$-uniform hypergraph is called an $\ell$-cycle if there exists a cyclic ordering of its vertices such that each of its edges consists of $r$ consecutive vertices, and such that every pair of…
The rate of convergence of the distribution of the length of the longest increasing subsequence, toward the maximal eigenvalue of certain matrix ensembles, is investigated. For finite-alphabet uniform and nonuniform i.i.d. sources, a rate…
We examine the Bayes-consistency of a recently proposed 1-nearest-neighbor-based multiclass learning algorithm. This algorithm is derived from sample compression bounds and enjoys the statistical advantages of tight, fully empirical…