Related papers: On (not) learning the M\"obius function
We present a unifying picture of PAC-Bayesian and mutual information-based upper bounds on the generalization error of randomized learning algorithms. As we show, Tong Zhang's information exponential inequality (IEI) gives a general recipe…
We consider polynomial equations, or systems of polynomial equations, with integer coefficients, modulo prime numbers $p$. We offer an elementary approach based on a counting method. The outcome is a weak form of the Lang-Weil lower bound…
Bayesian coresets speed up posterior inference in the large-scale data regime by approximating the full-data log-likelihood function with a surrogate log-likelihood based on a small, weighted subset of the data. But while Bayesian coresets…
In this work, we initiate the study of learning quantum processes from quantum statistical queries. We focus on two fundamental learning tasks in this new access model: shadow tomography of quantum processes and process tomography with…
Neural networks with binary weights are computation-efficient and hardware-friendly, but their training is challenging because it involves a discrete optimization problem. Surprisingly, ignoring the discrete nature of the problem and using…
In classical prime number theory there are several asymptotic formulas said to be "equivalent" to the PNT. One is the bound $M(x) = o(x)$ for the sum function of the Moebius function. For Beurling generalized numbers, this estimate is not…
While many approaches exist in the literature to learn low-dimensional representations for data collections in multiple modalities, the generalizability of multi-modal nonlinear embeddings to previously unseen data is a rather overlooked…
This paper proposes lower bounds on a quantity called $L^p$-norm joint spectral radius, or in short, $p$-radius, of a finite set of matrices. Despite its wide range of applications to, for example, stability analysis of switched linear…
The paper describes relations between Liouville type theorems for solutions of a periodic elliptic equation (or a system) on an abelian cover of a compact Riemannian manifold and the structure of the dispersion relation for this equation at…
We propose an adaptive zeroth-order method for minimizing differentiable functions with $L$-Lipschitz continuous gradients. The method is designed to take advantage of the eventual compressibility of the gradient of the objective function,…
Machine learning models trained by different optimization algorithms under different data distributions can exhibit distinct generalization behaviors. In this paper, we analyze the generalization of models trained by noisy iterative…
We present a semi-decision procedure to tackle first order differential equations, with Liouvillian functions in the solution (LFOODEs). As in the case of the Prelle-Singer procedure, this method is based on the knowledge of the integrating…
The choice of the kernel is critical to the success of many learning algorithms but it is typically left to the user. Instead, the training data can be used to learn the kernel by selecting it out of a given family, such as that of…
In this article we give several new results on the complexity of algorithms that learn Boolean functions from quantum queries and quantum examples. Hunziker et al. conjectured that for any class C of Boolean functions, the number of quantum…
We use a method of rotations to study the $L^p$ boundedness, $1<p<\infty$, of Fourier multipliers which arise as the projection of martingale transforms with respect to symmetric $\alpha$-stable processes, $0<\alpha<2$. Our proof does not…
The dictionary learning problem can be viewed as a data-driven process to learn a suitable transformation so that data is sparsely represented directly from example data. In this paper, we examine the problem of learning a dictionary that…
Aiming at optimizing the shape of closed embedded curves within prescribed isotopy classes, we use a gradient-based approach to approximate stationary points of the M\"obius energy. The gradients are computed with respect to Sobolev inner…
We study the problem of agnostic learning under the Gaussian distribution. We develop a method for finding hard families of examples for a wide class of problems by using LP duality. For Boolean-valued concept classes, we show that the…
This paper studies the generalization performance of multi-class classification algorithms, for which we obtain, for the first time, a data-dependent generalization error bound with a logarithmic dependence on the class size, substantially…
We are interested in classical and logarithmic imaginary classes of abelian number fields in connection with Iwasawa theory. For any given odd prime ${\ell}$ and any imaginary abelian number field K, we compute the isotypic components of…