Related papers: Arithmetic functions and learning theory
We provide examples of multiplicative functions $f$ supported on the squarefree integers, such that on primes $f(p)=\pm1$ and such that $M_f(x):=\sum_{n\leq x} f(n)=o(\sqrt{x})$. Further, by assuming the Riemann hypothesis (RH) we can go…
Generalization to out-of-distribution (OOD) data is one of the central problems in modern machine learning. Recently, there is a surge of attempts to propose algorithms that mainly build upon the idea of extracting invariant features.…
We study the task of bandit learning, also known as best-arm identification, under the assumption that the true reward function f belongs to a known, but arbitrary, function class F. We seek a general theory of bandit learnability, akin to…
We consider two classes of piecewise expanding maps $T$ of $[0,1]$: a class of uniformly expanding maps for which the Perron-Frobenius operator has a spectral gap in the space of bounded variation functions, and a class of expanding maps…
A classical problem in number theory is showing that the mean value of an arithmetic function is asymptotic to its mean value over a short interval or over an arithmetic progression, with the interval as short as possible or the modulus as…
We study the problem of approximating and learning coverage functions. A function $c: 2^{[n]} \rightarrow \mathbf{R}^{+}$ is a coverage function, if there exists a universe $U$ with non-negative weights $w(u)$ for each $u \in U$ and subsets…
We consider the classical problem of learning rates for classes with finite VC dimension. It is well known that fast learning rates up to $O\left(\frac{d}{n}\right)$ are achievable by the empirical risk minimization algorithm (ERM) if low…
We say that a classifier is \emph{adversarially robust} to perturbations of norm $r$ if, with high probability over a point $x$ drawn from the input distribution, there is no point within distance $\le r$ from $x$ that is classified…
Uniformity testing is arguably one of the most fundamental distribution testing problems. Given sample access to an unknown distribution $\mathbf{p}$ on $[n]$, one must decide if $\mathbf{p}$ is uniform or $\varepsilon$-far from uniform (in…
We prove that Sarnak's conjecture holds for any infinite measure symbolic rank-one map. We further extended Bourgain-Sarnak's result, which says that the M\"{o}bius function is a good weight for the ergodic theorem, to maps acting on…
We determine sufficient conditions under which certain recursively defined functions are well defined for all real inputs. Given a function $f:\mathbb R\to\mathbb R$, call a decreasing sequence $x_1>x_2>x_3>\cdots$ "$f$-bad" if…
We study upper bounds, approximations, and limits for functions of motivic exponential class, uniformly in non-Archimedean local fields whose characteristic is $0$ or sufficiently large. Our results together form a flexible framework for…
The model of relative-error property testing of Boolean functions has been the subject of significant recent research effort [CDH+24][CPPS25a][CPPS25b] In this paper we consider the problem of relative-error testing an unknown and arbitrary…
We present some tools for providing situations where the generalised Rota formula of arXiv:1801.07504 applies. As an example of this, we compute the M\"obius function of the incidence algebra of any directed restriction species, free…
We give a $2^{\tilde{O}(\sqrt{n}/\epsilon)}$-time algorithm for properly learning monotone Boolean functions under the uniform distribution over $\{0,1\}^n$. Our algorithm is robust to adversarial label noise and has a running time nearly…
Every computable function has to be continuous. To develop computability theory of discontinuous functions, we study low levels of the arithmetical hierarchy of nonuniformly computable functions on Baire space. First, we classify…
Partial observability is a common challenge in many reinforcement learning applications, which requires an agent to maintain memory, infer latent states, and integrate this past information into exploration. This challenge leads to a number…
The amount of training-data is one of the key factors which determines the generalization capacity of learning algorithms. Intuitively, one expects the error rate to decrease as the amount of training-data increases. Perhaps surprisingly,…
Considering a linearly ordered set, we introduce its symmetric version, and endow it with two operations extending supremum and infimum, so as to obtain an algebraic structure close to a commutative ring. We show that imposing symmetry…
Offline reinforcement learning seeks to utilize offline (observational) data to guide the learning of (causal) sequential decision making strategies. The hope is that offline reinforcement learning coupled with function approximation…