Related papers: Large Deviation Theory for Parameter Estimation in…
While suitably scaled CNNs with Gaussian initialization are known to converge to Gaussian processes as the number of channels diverges, little is known beyond this Gaussian limit. We establish a large deviation principle (LDP) for…
For the Ornstein-Uhlenbeck process, the asymptotic behavior of the maximum likelihood estimator of the drift parameter is totally different in the stable, unstable, and explosive cases. Notwithstanding of this trichotomy, we investigate…
We study the dynamics of on-line learning in large perceptrons, for the case of training sets with a structural bias of the input vectors, by deriving exact and closed macroscopic dynamical laws using non-equilibrium statistical mechanical…
Iterative preference optimization has recently become one of the de-facto training paradigms for large language models (LLMs), but the performance is still underwhelming due to too much noisy preference data yielded in the loop. To combat…
Noise-induced transitions between multistable states happen in a multitude of systems, such as species extinction in biology, protein folding, or tipping points in climate science. Large deviation theory is the rigorous language to describe…
Artificial neural networks can harness stochasticity in multiple ways to enable a vast class of computationally powerful models. Electronic implementation of such stochastic networks is currently limited to addition of algorithmic noise to…
In this paper, we propose a recurrent neural network (RNN)-based framework for estimating the parameters of the fractional Poisson process (FPP), which models event arrivals with memory and long-range dependence. The Long Short-Term Memory…
Latent variable models have been playing a central role in psychometrics and related fields. In many modern applications, the inference based on latent variable models involves one or several of the following features: (1) the presence of…
Assuming that a threshold Ornstein-Uhlenbeck process is observed at discrete time instants, we propose generalized moment estimators to estimate the parameters. Our theoretical basis is the celebrated ergodic theorem. To use this theorem we…
Mean field theory has been successfully used to analyze deep neural networks (DNN) in the infinite size limit. Given the finite size of realistic DNN, we utilize the large deviation theory and path integral analysis to study the deviation…
The fractional Ornstein-Uhleneck (fOU) process is described by the overdamped Langevin equation $\dot{x}(t)+\gamma x=\sqrt{2 D}\xi(t)$, where $\xi(t)$ is the fractional Gaussian noise with the Hurst exponent $0<H<1$. For $H\neq 1/2$ the fOU…
We develop here a stochastic framework for modeling and segmenting transient spindle-like oscillatory bursts in electroencephalogram (EEG) signals. At the modeling level, individual spindles are represented as path realizations of a…
We continue the development, started in of the asymptotic description of certain stochastic neural networks. We use the Large Deviation Principle (LDP) and the good rate function H announced there to prove that H has a unique minimum mu_e,…
For most stochastic dynamical systems, variables which are tightly regulated tend to respond slowly to external changes. This idea is often discussed for applicable systems, within a linear response regime, through the Fluctuation…
The spiking activity of single neurons can be well described by a nonlinear integrate-and-fire model that includes somatic adaptation. When exposed to fluctuating inputs sparsely coupled populations of these model neurons exhibit stochastic…
In this paper we develop a framework for estimating Probability of Default (PD) based on stochastic models governing an appropriate asset value processes. In particular, we build upon a L\'evy-driven Ornstein-Uhlenbeck process and consider…
The stochastic Hodgkin-Huxley neurons considered in this paper replace time-constant deterministic input $a dt$ of the classical deterministic model by increments $\vartheta dt + dX_t$ of a stochastic process: $X$ is Ornstein-Uhlenbeck with…
Large Language Models (LLMs) are composed of neurons that exhibit various behaviors and roles, which become increasingly diversified as models scale. Recent studies have revealed that not all neurons are active across different datasets,…
Learning is a fundamental property of intelligent systems, observed across biological organisms and engineered systems. While modern intelligent systems typically rely on gradient descent for learning, the need for exact gradients and…
The theory of large deviations constitutes a mathematical cornerstone in the foundations of Boltzmann-Gibbs statistical mechanics, based on the additive entropy $S_{BG}=- k_B\sum_{i=1}^W p_i \ln p_i$. Its optimization under appropriate…