Related papers: Attention in Krylov Space
Graph-based next-step prediction models have recently been very successful in modeling complex high-dimensional physical systems on irregular meshes. However, due to their short temporal attention span, these models suffer from error…
In this paper, we studied a set of generalised Krylov complexity for operator growth. We demonstrate their universal features at both initial times and long times using half-analytical technique as well as numerical results. In particular,…
This paper introduces a high-order Markov chain task to investigate how transformers learn to integrate information from multiple past positions with varying statistical significance. We demonstrate that transformers learn this task…
In this paper we investigate the long time behavior of solutions to fractional in time evolution equations which appear as results of random time changes in Markov processes. We consider inverse subordinators as random times and use the…
A persistent paradox in time-series forecasting is that structurally simple MLP and linear models often outperform high-capacity Transformers. We argue that this gap arises from a mismatch in the sequence-modeling primitive: while many…
We investigate the complexity of states and operators evolved with the modular Hamiltonian by using the Krylov basis. In the first part, we formulate the problem for states and analyse different examples, including quantum mechanics,…
The Krylov subspace method is a standard approach to approximate quantum evolution, allowing to treat systems with large Hilbert spaces. Although its application is general, and suitable for many-body systems, estimation of the committed…
The transformer is the most popular neural architecture for language modeling. The cornerstone of the transformer is its global attention mechanism, which lets the model aggregate information from all preceding tokens before generating the…
Modern language models rely on the transformer architecture and attention mechanism to perform language understanding and text generation. In this work, we study learning a 1-layer self-attention model from a set of prompts and associated…
Scrambling of information in a quantum many-body system, quantified by the out-of-time-ordered correlator (OTOC), is a key manifestation of quantum chaos. A regime of exponential growth in the OTOC, characterized by a Lyapunov exponent, has…
Transformers can under some circumstances generalize to novel problem instances whose constituent parts might have been encountered during training, but whose compositions have not. What mechanisms underlie this ability for compositional…
Since its introduction, the transformer has shifted the development trajectory away from traditional models (e.g., RNN, MLP) in time series forecasting, which is attributed to its ability to capture global dependencies within temporal…
Large language models (LLMs) based on transformer architectures are typically described through collections of architectural components and training procedures, obscuring their underlying computational structure. This review article…
Phase transitions mark qualitative reorganizations of collective behavior, yet identifying their boundaries remains challenging whenever analytic solutions are absent and conventional simulations fail. Here we introduce learnability as a…
We study the evolution of distributions under the action of an ergodic dynamical system, which may be stochastic in nature. By employing tools from Koopman and transfer operator theory one can evolve any initial distribution of the state…
The predictability problem for systems with different characteristic time scales is investigated. It is shown that even in simple chaotic dynamical systems, the leading Lyapunov exponent is not sufficient to estimate the predictability…
Classical quasi-integrable systems are known to have Lyapunov times much shorter than their ergodicity time, but the situation for their quantum counterparts is less well understood. As a first example, we examine the quantum Lyapunov…
We reformulate the Lanczos algorithm for quantum wave function propagation in terms of variational principle. By including some basis states of previous time steps into the variational subspace, the resultant accuracy increases by several…
This paper explores the asymptotic behavior of univariate neural network operators, with an emphasis on both classical and fractional differentiation over infinite domains. The analysis leverages symmetrized and perturbed hyperbolic tangent…
The use of machine learning methods for predictive purposes has increased dramatically over the past two decades, but uncertainty quantification for predictive comparisons remains elusive. This paper addresses this gap by extending the…