Related papers: Vertex-reinforced jump process on the integers wit…
Questions remain on the robustness of data-driven learning methods when crossing the gap from simulation to reality. We utilize weight anchoring, a method known from continual learning, to cultivate and fixate desired behavior in Neural…
We investigate the possibility to determine the divergence-free displacement $\mathbf{u}$ \emph{independently} from the pressure reaction $p$ for a class of boundary value problems in incompressible linear elasticity. If not possible, we…
This paper is devoted to studying an infinite time horizon stochastic recursive control problem with jumps, where infinite time horizon stochastic differential equation and backward stochastic differential equation with jumps describe the…
We introduce Leap+Verify, a framework that applies speculative execution -- predicting future model weights and validating predictions before acceptance -- to accelerate neural network training. Inspired by speculative decoding in language…
The point process of vertices of an iteration infinitely divisible or more specifically of an iteration stable random tessellation in the Euclidean plane is considered. We explicitly determine its covariance measure and its pair-correlation…
A $\delta$ once-reinforced random walk ($\delta$-ORRW) on connected graph is a self-interacting random walk which moves to its neighbors at each step according to the weights of the edges at that time, where the weights are $1$ on edges…
Reinforcement Learning with Verifiable Rewards (RLVR) has proven effective for enhancing Large Language Models (LLMs) on complex reasoning tasks. However, existing methods suffer from an exploration dilemma: the sharply peaked initial…
An artificial neural network can be trained by uniformly broadcasting a reward signal to units that implement a REINFORCE learning rule. Though this presents a biologically plausible alternative to backpropagation in training a network, the…
Reinforcement learning (RL) has shown great potential in enabling quadruped robots to perform agile locomotion. However, directly training policies to simultaneously handle dual extreme challenges, i.e., extreme underactuation and extreme…
In this paper, we are interested in the asymptotic behaviour of the sequence of processes $(W_n(s,t))_{s,t\in[0,1]}$ with \begin{equation*} W_n(s,t):=\sum_{k=1}^{\lfloor nt\rfloor}\big(1_{\{\xi_{S_k}\leq s\}}-s\big) \end{equation*} where…
We study perturbatively the (conformal) WZNW model. At one loop we compute one-particle irreducible two- and three-point current correlation functions, both in the conventional version and in the classically equivalent, chiral, nonlocal,…
We consider the interlacement Poisson point process on the space of doubly-infinite Z^d-valued trajectories modulo time-shift, tending to infinity at positive and negative infinite times. The set of vertices and edges visited by at least…
Based on a martingale theory approach, we present a complete characterization of the asymptotic behaviour of a lazy reinforced random walk (LRRW) which shows three different regimes (diffusive, critical and superdiffusive). This allows us…
Consider the random set composed of particles initially distributed on Zd, d >= 2, according to a Poisson point process of intensity u > 0 and moving as independent simple symmetric random walks, the trap particles. We are interested in the…
This work deals with systems of interacting reinforced stochastic processes, where each process $X^j=(X_{n,j})_n$ is located at a vertex $j$ of a finite weighted direct graph, and it can be interpreted as the sequence of "actions" adopted…
As deep neural networks are increasingly being deployed in practice, their efficiency has become an important issue. While there are compression techniques for reducing the network's size, energy consumption and computational requirement,…
In recent studies on sparse modeling, non-convex penalties have received considerable attentions due to their superiorities on sparsity-inducing over the convex counterparts. Compared with the convex optimization approaches, however, the…
We investigate the dynamics of a particle executing a general Continuous Time Random Walk (CTRW) in three dimensions under the influence of arbitrary time-varying external fields. Contrary to the general approach in recent works, our method…
By extracting unstable invariant solutions directly from body-forced three-dimensional turbulence, we study the dynamical processes at play when the forcing is large scale and either unidirectional in the momentum or the vorticity…
We study the problem of infinite-horizon average-reward reinforcement learning with linear Markov decision processes (MDPs). The associated Bellman operator of the problem not being a contraction makes the algorithm design challenging.…