Related papers: How to explain grokking
We consider the problem of building a continuous stochastic model, i.e. a Langevin or Fokker-Planck equation, through a well-controlled coarse-graining procedure. Such a method usually involves the elimination of the fast degrees of freedom…
We analyze some basic issues associated with Generalized Poisson-Kac (GPK) stochastic processes, starting from the extended notion of the Markovian condition. The extended Markovian nature of GPK processes is established, and the…
We investigate stochastic processes that generalize geometric Brownian motion, focusing on cases where the standard invariant measure, i.e. the solution of the stationary Fokker-Planck equation does not necessarily exist. We demonstrate…
Language provides simple ways of communicating generalizable knowledge to each other (e.g., "Birds fly", "John hikes", "Fire makes smoke"). Though found in every language and emerging early in development, the language of generalization is…
A short review of the classical theory of Brownian motion is presented. A new method is proposed for derivation of the Fokker-Planck equations, describing the probability density evolution, from stochastic differential equations. It is also…
This is a guide to the mathematical theory of Brownian motion and related stochastic processes, with indications of how this theory is related to other branches of mathematics, most notably the classical theory of partial differential…
Training dynamics during grokking concentrate along a small number of dominant update directions -- the spectral edge -- which reliably distinguishes grokking from non-grokking regimes. We show that standard mechanistic interpretability…
Physical scenarios that require a relativistic treatment are ubiquitous in nature, ranging from cosmological objects to charge carriers in Dirac materials. Interestingly all of these situations have in common that the systems typically…
The generalised Langevin equation with a retarded friction and a double-well potential is solved. The random force is modelled by a multiplicative noise with long jumps. Probability density distributions converge with time to a distribution…
The focus of our study in this paper is on the active dynamics and a fractional generalized Langevin equation with a memory kernel K(t). The Fokker-Planck equation is obtained by deriving it from a second-order differential equation. The…
Recently there has been remarkable progress in the complex Langevin method, which aims at solving the complex action problem by complexifying the dynamical variables in the original path integral. In particular, a new technique called the…
This paper investigates the impact of different optimizers on the grokking phenomenon, where models exhibit delayed generalization. We conducted experiments across seven numerical tasks (primarily modular arithmetic) using a modern…
A model has two main aims: predicting the behavior of a physical system and understanding its nature, that is how it works, at some desired level of abstraction. A promising recent approach to model building consists in deriving a…
Stochastic inflation rests on the separate-universe approximation, i.e. the ability to describe long-wavelength fluctuations in an inflating universe as homogeneous perturbations of its background dynamics. Although this approximation is…
Nonergodic Brownian motion is elucidated within the framework of the generalized Langevin equation. For thermal noise yielding either a vanishing or a divergent zero-frequency friction strength, the non-Markovian Browninan dynamics exhibits…
The fluctuation-dissipation theorem is a central result in statistical mechanics and is usually formulated for systems described by diffusion processes. In this paper, we propose a generalization for a wider class of stochastic processes,…
We present a numerical method to produce stochastic dynamics according to the generalized Langevin equation with a non-stationary memory kernel. This type of dynamics occurs when a microscopic system with an explicitly time-dependent…
We design and analyze a new paradigm for building supervised learning networks, driven only by local optimization rules without relying on a global error function. Traditional neural networks with a fixed topology are made up of identical…
The coarse-graining approach to deriving the quantum Markovian master equation is revisited, with close attention given to the underlying approximations. It is further argued that the time interval over which the coarse-graining is…
A major challenge in reinforcement learning is to determine which state-action pairs are responsible for future rewards that are delayed. Reward redistribution serves as a solution to re-assign credits for each time step from observed…