Related papers: Approximation of hyperarithmetic analysis by $\ome…
Reinforcement learning from human feedback (RLHF) replaces hard-to-specify rewards with pairwise trajectory preferences, yet regret-oriented theory often assumes that preference labels are generated consistently from a single ground-truth…
Quantum affine reflection algebras are coideal subalgebras of quantum affine algebras that lead to trigonometric reflection matrices (solutions of the boundary Yang-Baxter equation). In this paper we use the quantum affine reflection…
Building systems that autonomously create temporal abstractions from data is a key challenge in scaling learning and planning in reinforcement learning. One popular approach for addressing this challenge is the options framework (Sutton et…
In this note we give a simplified ordinal analysis of first-order reflection. An ordinal notation system $OT$ is introduced based on $\psi$-functions. Provable $\Sigma_{1}$-sentences on $L_{\omega_{1}^{CK}}$ are bounded through…
We show that when certain statements are provable in subsystems of constructive analysis using intuitionistic predicate calculus, related sequential statements are provable in weak classical subsystems. In particular, if a $\Pi^1_2$…
Abstraction reasoning is a long-standing challenge in artificial intelligence. Recent studies suggest that many of the deep architectures that have triumphed over other domains failed to work well in abstract reasoning. In this paper, we…
We consider extensions of the language of Peano arithmetic by transfinitely iterated truth definitions satisfying uniform Tarskian biconditionals. Without further axioms, such theories are known to be conservative extensions of the original…
Reinforcement learning with outcome-based feedback faces a fundamental challenge: when rewards are only observed at trajectory endpoints, how do we assign credit to the right actions? This paper provides the first comprehensive analysis of…
Backpropagation is driving today's artificial neural networks (ANNs). However, despite extensive research, it remains unclear if the brain implements this algorithm. Among neuroscientists, reinforcement learning (RL) algorithms are often…
The theory of reinforcement learning has focused on two fundamental problems: achieving low regret, and identifying $\epsilon$-optimal policies. While a simple reduction allows one to apply a low-regret algorithm to obtain an…
We study a quadruple of interrelated subexponential subsystems of arithmetic WKL$_0^-$, RCA$^-_0$, I$\Delta_0$, and $\Delta$RA$_1$, which complement the similarly related quadruple WKL$_0$, RCA$_0$, I$\Sigma_1$, and PRA studied by Simpson,…
This paper studies three results that describe the structure of the super-coinvariant algebra of pseudo-reflection groups over a field of characteristic $0$. Our most general result determines the top component in total degree, which we…
In this paper, we introduce a hierarchy dividing the set $\{\sigma \in \Pi^1_2 : \Pi^1_1$-$\mathsf{CA}_0 \vdash \sigma\}$. Then, we give some characterizations of this set using weaker variants of some principles equivalent to…
We introduce two approximate variants of inclusion dependencies and examine the axiomatization and computational complexity of their implication problems. The approximate variants allow for some imperfection in the database and differ in…
We deal with the \textit{selective classification} problem (supervised-learning problem with a rejection option), where we want to achieve the best performance at a certain level of coverage of the data. We transform the original $m$-class…
This paper is a sequel to our earlier work [BFPRW], where we study the derived representation scheme DRep_{g}(A) parametrizing the representations of a Lie algebra A in a finite-dimensional reductive Lie algebra g. In [BFPRW], we defined…
We develop a behavioural theory of reflective sequential algorithms (RSAs), i.e. sequential algorithms that can modify their own behaviour. The theory comprises a set of language-independent postulates defining the class of RSAs, an…
We consider increasingly complex models of matrix denoising and dictionary learning in the Bayes-optimal setting, in the challenging regime where the matrices to infer have a rank growing linearly with the system size. This is in contrast…
Model-free reinforcement learning algorithms combined with value function approximation have recently achieved impressive performance in a variety of application domains. However, the theoretical understanding of such algorithms is limited,…
Given a database and a target attribute of interest, how can we tell whether there exists a functional, or approximately functional dependence of the target on any set of other attributes in the data? How can we reliably, without bias to…