Related papers: Learning low-degree functions from a logarithmic n…
For $V : \mathbb{R}^d \to \mathbb{R}$ coercive, we study the convergence rate for the $L^1$-distance of the empiric minimizer, which is the true minimum of the function $V$ sampled with noise with a finite number $n$ of samples, to the…
A loss function measures the discrepancy between the true values (observations) and their estimated fits, for a given instance of data. A loss function is said to be proper (unbiased, Fisher consistent) if the fits are defined over a unit…
We prove a Lusin approximation of functions of bounded variation. If $f$ is a function of bounded variation on an open set $\Omega\subset X$, where $X=(X,d,\mu)$ is a given complete doubling metric measure space supporting a $1$-Poincar\'e…
Let $X$ be a real-valued random variable with distribution function $F$. Set $X_1,\dots, X_m$ to be independent copies of $X$ and let $F_m$ be the corresponding empirical distribution function. We show that there are absolute constants…
Let $\Lambda = \mathrm{SL}_2(\Bbb Z)$ be the modular group and let $c_n(\Lambda)$ be the number of congruence subgroups of $\Lambda$ of index at most $n$. We prove that $\lim\limits_{n\to \infty} \frac{\log c_n(\Lambda)}{(\log n)^2/\log\log…
Loss functions serve as the foundation of supervised learning and are often chosen prior to model development. To avoid potentially ad hoc choices of losses, statistical decision theory describes a desirable property for losses known as…
The Local Computation Algorithm (LCA) model is a popular model in the field of sublinear-time algorithms that measures the complexity of an algorithm by the number of probes the algorithm makes in the neighborhood of one node to determine…
Littlewood asked how small the ratio $||f||_4/||f||_2$ (where $||.||_\alpha$ denotes the $L^\alpha$ norm on the unit circle) can be for polynomials $f$ having all coefficients in $\{1,-1\}$, as the degree tends to infinity. Since 1988, the…
The theorem states that: Every Boolean function can be $\epsilon -approximated$ by a Disjunctive Normal Form (DNF) of size $O_{\epsilon}(2^{n}/\log{n})$. This paper will demonstrate this theorem in detail by showing how this theorem is…
A classical theorem of Nisan and Szegedy says that a boolean function with degree $d$ as a real polynomial depends on at most $d2^{d-1}$ of its variables. In recent work by Chiarelli, Hatami and Saks, this upper bound was improved to $C…
We study batch learning with log-loss in the individual setting, where the outcome sequence is deterministic. Because empirical statistics are not directly applicable in this regime, obtaining regret guarantees for batch learning has long…
We prove that the finite-difference based derivative-free descent (FD-DFD) methods have a capability to find the global minima for a class of multiple minima problems. Our main result shows that, for a class of multiple minima objectives…
We study the approximation of measurable functions on the hypercube by functions arising from affine neural networks. Our main achievement is an approximation of any measurable function $f \colon W_n \to [-1,1]$ up to a prescribed precision…
We present an algorithm for testing halfspaces over arbitrary, unknown rotation-invariant distributions. Using $\tilde O(\sqrt{n}\epsilon^{-7})$ random examples of an unknown function $f$, the algorithm determines with high probability…
We present effective algorithms for uniform approximation of multivariate functions satisfying some prescribed inner structure. We extend in several directions the analysis of recovery of ridge functions $f(x)=g(\langle a,x\rangle)$ as…
This is a conitunation of [1] and [2]. We prove that if function $f$ belongs to the class $\Lambda_{\omega} \overset{\text{def}}{=} \{f: \omega_{f}(\delta)\leq \text{const} \omega(\delta)\} $ for an arbitrary modulus of continuity $\omega$,…
In this paper, we study the learnability of the Boolean class of $d$-monotone functions $f:{\cal X}\to\{0,1\}$ from membership and equivalence queries, where $({\cal X},\le)$ is a finite lattice. We show that the class of $d$-monotone…
We give a lower bound for the maximum value of class group $L$-functions attached to $\mathbb{Q}(\sqrt{-D})$ at the central point and show that this value is on average at least $$\exp\Bigg(\delta\sqrt{\frac{\log D \log \log \log D}{\log…
We prove that for every decision tree, the absolute values of the Fourier coefficients of a given order $\ell\geq1$ sum to at most $c^{\ell}\sqrt{\binom{d}{\ell}(1+\log n)^{\ell-1}},$ where $n$ is the number of variables, $d$ is the tree…
We show a procedure that, given oracle access to a function $f\colon \{0,1\}^n\to\{0,1\}$, produces oracle access to a function $f'\colon \{0,1\}^{n'}\to\{0,1\}$ such that if $f$ is monotone, then $f'$ is monotone, and if $f$ is…