Related papers: Vertex-reinforced jump process on the integers wit…
We consider a class of reinforcement processes, called WARMs, on tree graphs. These processes involve a parameter $\alpha$ which governs the strength of the reinforcement, and a collection of Poisson processes indexed by the vertices of the…
We investigate Bernoulli free boundary problems prescribing infinite jump conditions. The mathematical set-up leads to the analysis of non-differentiable minimization problems of the form $\int \left(\nabla u\cdot (A(x)\nabla u) +…
We address the challenge of exploration in reinforcement learning (RL) when the agent operates in an unknown environment with sparse or no rewards. In this work, we study the maximum entropy exploration problem of two different types. The…
Uniformly valid inference for cointegrated vector autoregressive processes has so far proven difficult due to certain discontinuities arising in the asymptotic distribution of the least squares estimator. We extend asymptotic results from…
Animals ranging from rats to humans can demonstrate cognitive map capabilities. We evolved weights in a biologically plausible recurrent neural network (RNN) using an evolutionary algorithm to replicate the behavior and neural activity…
We consider the activated random walk (ARW) model where particles follow the path of a general Markov process on a general graph. We prove ARW dominates a simpler process, multiple source internal aggregation (MSIA), and use this to…
This paper studies the non-implosion mechanism for the 3D incompressible Euler equations. We prove that vorticity blows up in finite time, whereas the $L^p_T L^\infty_{loc}$ $(p\in[1,\infty))$ norm of the velocity field remains bounded.…
In this study, we investigate the continuous time dynamics of Recurrent Neural Networks (RNNs), focusing on systems with nonlinear activation functions. The objective of this work is to identify conditions under which RNNs exhibit perpetual…
This paper reformulates complementarity-based time-stepping for frictionless nonsmooth contact between smooth rigid bodies as a recursively generated linear complementarity problem (ReLCP), involving a sequence of LCPs of increasing…
Algebraic random walks (ARW) and quantum mechanical random walks (QRW) are investigated and related. Based on minimal data provided by the underlying bialgebras of functions defined on e. g the real line R, the abelian finite group Z_N, and…
Let $\boldsymbol W=\{\boldsymbol W_n:n\in\mathbb N\}$ be a sequence of random vectors in $\mathbb R^d$, $d\ge 1$. This paper considers the logarithmic asymptotics of the extremes of $\boldsymbol W$, that is, for any vector $\boldsymbol…
We consider un-discounted reinforcement learning (RL) in Markov decision processes (MDPs) under drifting non-stationarity, i.e., both the reward and state transition distributions are allowed to evolve over time, as long as their respective…
We consider the Reinforcement Learning problem of controlling an unknown dynamical system to maximise the long-term average reward along a single trajectory. Most of the literature considers system interactions that occur in discrete time…
Reinforcement learning from verifiable rewards (RLVR) has improved the reasoning abilities of large language models (LLMs), but most existing approaches rely on sparse outcome-level feedback. This sparsity creates a credit assignment…
We consider activated random walk (ARW), an interacting particle system and prototypical model of self-organized criticality in a setting which combines mean-field behavior with the geometry of an arbitrary graph, which we call the village…
We show that reinforcement learning with verifiable rewards (RLVR) can elicit strong mathematical reasoning in certain language models even with spurious rewards that have little, no, or even negative correlation with the correct answer.…
Reinforcement Learning with Verifiable Rewards (RLVR) serves as a cornerstone technique for enhancing the reasoning capabilities of Large Language Models (LLMs). However, its training is often plagued by \emph{entropy collapse}, a rapid…
Reconstruction of signals from undersampled and noisy measurements is a topic of considerable interest. Sharpness conditions directly control the recovery performance of restart schemes for first-order methods without the need for…
The Weighted Vertex Integrity (wVI) problem takes as input an $n$-vertex graph $G$, a weight function $w:V(G)\to\mathbb{N}$, and an integer $p$. The task is to decide if there exists a set $X\subseteq V(G)$ such that the weight of $X$ plus…
Fast weight architectures offer a promising alternative to attention-based transformers for long-context modeling by maintaining constant memory overhead regardless of context length. However, their potential is limited by the next-token…