Related papers: The height function of a sparse collection: a Bell…
While sparse attention mitigates the computational bottleneck of long-context LLM training, its distributed training process exhibits extreme heterogeneity in both \textit{1)} sequence length and \textit{2)} sparsity sensitivity, leading to…
This paper proposes a new method for estimating sparse precision matrices in the high dimensional setting. It has been popular to study fast computation and adaptive procedures for this problem. We propose a novel approach, called Sparse…
We provide some new estimates for Bellman type functions for the dyadic maximal opeator on $R^n$ and of maximal operators on martingales related to weighted spaces. Using a type of symmetrization principle, introduced for the dyadic maximal…
The spectrum and coherency are useful quantities for characterizing the temporal correlations and functional relations within and between point processes. This paper begins with a review of these quantities, their interpretation and how…
Sparse representation has been applied successfully in abnormal event detection, in which the baseline is to learn a dictionary accompanied by sparse codes. While much emphasis is put on discriminative dictionary construction, there are no…
In this paper, we study pointwise estimates for linear and multilinear pseudo-differential operators with exotic symbols in terms of the Fefferman-Stein sharp maximal function and Hardy-Littlewood type maximal function. Especially in the…
In sparse Bayesian learning (SBL), Gaussian scale mixtures (GSMs) have been used to model sparsity-inducing priors that realize a class of concave penalty functions for the regression task in real-valued signal models. Motivated by the…
We get sharp estimates for the distribution function of nonnegative weights, which satisfy so called $A_{p_1, p_2}$ condition. For particular choices of parameters $p_1$, $p_2$ this condition becomes an $A_p$-condition or Reverse H\"{o}lder…
We propose sparsemax, a new activation function similar to the traditional softmax, but able to output sparse probabilities. After deriving its properties, we show how its Jacobian can be efficiently computed, enabling its use in a network…
Background: High-throughput proteomics techniques, such as mass spectrometry (MS)-based approaches, produce very high-dimensional data-sets. In a clinical setting one is often interested in how mass spectra differ between patients of…
We consider interpolation inequalities for imbeddings of the $l^2$-sequence spaces over $d$-dimensional lattices into the $l^\infty_0$ spaces written as interpolation inequality between the $l^2$-norm of a sequence and its difference. A…
A symmetric function of $N$ variables can be given in terms of symmetric polynomials of these variables. We determine those symmetric polynomials in which the dual differential operators take the neatest form when expressed in terms of our…
Sparse optimization is a central problem in machine learning and computer vision. However, this problem is inherently NP-hard and thus difficult to solve in general. Combinatorial search methods find the global optimal solution but are…
We find sharp constants in the symmetric integral form of the John-Nirenberg inequality. The result is based upon computation of a new interesting Bellman function.
The problem of consistently estimating the sparsity pattern of a vector $\betastar \in \real^\mdim$ based on observations contaminated by noise arises in various contexts, including subset selection in regression, structure estimation in…
For a class of sparse operators including majorants of singular integral, square function, and fractional integral operators in a uniform manner, we prove off-diagonal two-weight estimates of mixed type in the two-weight and…
Approximation of high-dimensional functions is a problem in many scientific fields that is only feasible if advantageous structural properties, such as sparsity in a given basis, can be exploited. A relevant tool for analysing sparse…
We consider the Ensemble Kalman Inversion which has been recently introduced as an efficient, gradient-free optimisation method to estimate unknown parameters in an inverse setting. In the case of large data sets, the Ensemble Kalman…
Sparse prediction with categorical data is challenging even for a moderate number of variables, because one parameter is roughly needed to encode one category or level. The Group Lasso is a well known efficient algorithm for selection…
Neural recordings, returns from radars and sonars, images in astronomy and single-molecule microscopy can be modeled as a linear superposition of a small number of scaled and delayed copies of a band-limited or diffraction-limited point…