Related papers: Arithmetic functions and learning theory
We consider a large class of piecewise expanding maps T of [0,1] with a neutral fixed point, and their associated Markov chain Y_i whose transition kernel is the Perron-Frobenius operator of T with respect to the absolutely continuous…
We show that $$ \sum_{n\neq m}\frac{\mu(n)\mu(m)}{nm}E_{X}\left(\{nx\}\{mx\}\right)=-\frac{9}{2\pi^{2}}+O\left(\frac{1}{X}\right), $$ where $x$ is uniformly distributed in $[0,X]$ with $X\in \mathbb{N}$, $E_{X}(.)$ denotes the expected…
Operator learning is a data-driven approximation of mappings between infinite-dimensional function spaces, such as the solution operators of partial differential equations. Kernel-based operator learning can offer accurate, theoretically…
We develop a general framework for estimating function-valued parameters under equality or inequality constraints in infinite-dimensional statistical models. Such constrained learning problems are common across many areas of statistics and…
We discuss the multiplicity of the non-trivial zeros of the Riemann zeta-function and the summatory function $M(x)$ of the M\"obius function. The purpose of this paper is to consider two open problems under some conjectures. One is that…
We consider the problem of learning an unknown $f$ with a sparse Fourier spectrum in the presence of outlier noise. In particular, the algorithm has access to a noisy oracle for (an unknown) $f$ such that (i) the Fourier spectrum of $f$ is…
The Mertens function is defined as $M(x) = \sum_{n \leq x} \mu(n)$, where $\mu(n)$ is the M\"obius function. The Mertens conjecture states $|M(x)/\sqrt{x}| < 1$ for $x > 1$, which was proven false in 1985 by showing $\liminf M(x)/\sqrt{x} <…
For every natural number k we prove a decomposition theorem for bounded measurable functions on compact abelian groups into a structured part, a quasi random part and a small error term. In this theorem quasi randomness is measured with the…
We consider the basic problem of learning Single-Index Models with respect to the square loss under the Gaussian distribution in the presence of adversarial label noise. Our main contribution is the first computationally efficient algorithm…
This article studies the achievable guarantees on the error rates of certain learning algorithms, with particular focus on refining logarithmic factors. Many of the results are based on a general technique for obtaining bounds on the error…
The problem of learning Boolean linear functions from quantum examples w.r.t. the uniform distribution can be solved on a quantum computer using the Bernstein-Vazirani algorithm. A similar strategy can be applied in the case of noisy…
The exploration-exploitation dilemma has been a central challenge in reinforcement learning (RL) with complex model classes. In this paper, we propose a new algorithm, Monotonic Q-Learning with Upper Confidence Bound (MQL-UCB) for RL with…
In this paper we give a short interval version of the Balog-Ruzsa theorem concerning bounds for the $L_1$ norm of the exponential sum over $r$-free numbers. As an application, we give a lower bound for the $L_1$ norm of the exponential sum…
We establish quantitative bounds on the $U^k[N]$ Gowers norms of the M\"obius function $\mu$ and the von Mangoldt function $\Lambda$ for all $k$, with error terms of shape $O((\log\log N)^{-c})$. As a consequence, we obtain quantitative…
We consider the variance of sums of arithmetic functions over random short intervals in the function field setting. Based on the analogy between factorizations of random elements of $\mathbb{F}_q[T]$ into primes and the factorizations of…
Most existing literature on supervised machine learning assumes that the training dataset is drawn from an i.i.d. sample. However, many real-world problems exhibit temporal dependence and strong correlations between the marginal…
M${}^{\natural}$-concave functions, a.k.a. gross substitute valuation functions, play a fundamental role in many fields, including discrete mathematics and economics. In practice, perfect knowledge of M${}^{\natural}$-concave functions is…
We find a sharp combinatorial bound for the metric entropy of sets in R^n and general classes of functions. This solves two basic combinatorial conjectures on the empirical processes. 1. A class of functions satisfies the uniform Central…
We develop minimax optimal risk bounds for the general learning task consisting in predicting as well as the best function in a reference set $\mathcal{G}$ up to the smallest possible additive term, called the convergence rate. When the…
We study a class of nonconvex nonsmooth optimization problems in which the objective is a sum of two functions: One function is the average of a large number of differentiable functions, while the other function is proper, lower…