English
Related papers

Related papers: Foundations and Fundamental Properties of a Two-Pa…

200 papers

Moving from univariate to bivariate jointly dependent long-memory time series introduces a phase parameter $(\gamma)$, at the frequency of principal interest, zero; for short-memory series $\gamma=0$ automatically. The latter case has also…

Statistics Theory · Mathematics 2008-11-07 P. M. Robinson

In model-based reinforcement learning, the transition matrix and reward vector are often estimated from random samples subject to noise. Even if the estimated model is an unbiased estimate of the true underlying model, the value function…

Machine Learning · Computer Science 2023-02-09 Xun Tang , Lexing Ying , Yuhua Zhu

Forgetting is an important concept in knowledge representation and automated reasoning with widespread applications across a number of disciplines. A standard forgetting operator, characterized in [Lin and Reiter'94] in terms of…

Artificial Intelligence · Computer Science 2024-12-06 Patrick Doherty , Andrzej Szalas

A possible solution for the problem of memory-size and computer-time, is the extrapolation of basis-set$^1$. This extrapolation has two exponents $\alpha$ and $\beta$, corresponding to the HF (reference energy) and the energy of…

Chemical Physics · Physics 2015-08-03 Suresh Chandra , Mohit K. Sharma

Memory and forgetting constitute two sides of the same coin, and although the first has been rigorously investigated, the latter is often overlooked. A number of experiments under the realm of psychology and experimental neuroscience have…

Neurons and Cognition · Quantitative Biology 2019-07-23 Antonios Georgiou , Mikhail Katkov , Misha Tsodyks

Use of generative models and deep learning for physics-based systems is currently dominated by the task of emulation. However, the remarkable flexibility offered by data-driven architectures would suggest to extend this representation to…

Machine Learning · Computer Science 2023-09-12 Guoxiang Grayson Tong , Carlos A. Sing Long , Daniele E. Schiavazzi

For over a decade, explicit memory architectures like the Neural Turing Machine have remained theoretically appealing yet practically intractable for language modeling due to catastrophic gradient instability during Backpropagation Through…

Machine Learning · Computer Science 2026-05-14 Sungwoo Goo , Hwi-yeol Yun , Sangkeun Jung

This paper establishes a rigorous spectral framework for the Weighted Weyl Fractional Calculus, designed to model non-local systems exhibiting aging and subjective time scales. By constructing a conjugation map involving a time-dependent…

Spectral Theory · Mathematics 2026-01-06 Gustavo Dorrego

Vision-language-action (VLA) models for closed-loop robot control are typically cast under the Markov assumption, making them prone to errors on tasks requiring historical context. To incorporate memory, existing VLAs either retrieve from a…

Robotics · Computer Science 2026-03-16 Hang Li , Fengyi Shen , Dong Chen , Liudi Yang , Xudong Wang , Jinkui Shi , Zhenshan Bing , Ziyuan Liu , Alois Knoll

We show the formal equivalence of linearised self-attention mechanisms and fast weight controllers from the early '90s, where a ``slow" neural net learns by gradient descent to program the ``fast weights" of another net through sequences of…

Machine Learning · Computer Science 2021-06-10 Imanol Schlag , Kazuki Irie , Jürgen Schmidhuber

Standard Transformers impose near-exponential decay on the influence of distant tokens, conflicting with the power-law structure of long-range dependencies in natural language. We introduce the \emph{Variable-Order Retention Transformer}…

Machine Learning · Computer Science 2026-05-12 Nabil Mlaiki

We prove limit theorems of an entirely new type for certain long memory regularly varying stationary infinitely divisible random processes. These theorems involve multiple phase transitions governed by how long the memory is. Apart from one…

Probability · Mathematics 2018-05-23 Gennady Samorodnitsky , Yizao Wang

The notion of disentangled autoencoders was proposed as an extension to the variational autoencoder by introducing a disentanglement parameter $\beta$, controlling the learning pressure put on the possible underlying latent representations.…

Machine Learning · Statistics 2017-11-28 Momchil Peychev , Petar Veličković , Pietro Liò

To accommodate structured approaches of neural computation, we propose a class of recurrent neural networks for indexing and storing sequences of symbols or analog data vectors. These networks with randomized input weights and orthogonal…

Neural and Evolutionary Computing · Computer Science 2018-03-02 E. Paxon Frady , Denis Kleyko , Friedrich T. Sommer

The time periodic circuit theory is exploited to introduce an appropriate translation operator that is invariant under the change of the spatial unit cell. Useful properties of the operator are derived. By casting the problem in an…

Applied Physics · Physics 2020-08-25 Sameh Y. Elnaggar , Gregory. N. Milford

In many data analysis tasks, it is beneficial to learn representations where each dimension is statistically independent and thus disentangled from the others. If data generating factors are also statistically independent, disentangled…

Machine Learning · Statistics 2019-12-12 Harshvardhan Sikka , Weishun Zhong , Jun Yin , Cengiz Pehlevan

Data-driven reduced-order models based on autoencoders generally lack interpretability compared to classical methods such as the proper orthogonal decomposition. More interpretability can be gained by disentangling the latent variables and…

Machine Learning · Computer Science 2025-02-21 Henning Schwarz , Pyei Phyo Lin , Jens-Peter M. Zemke , Thomas Rung

Variational autoencoders (VAEs) are powerful tools for learning latent representations of data used in a wide range of applications. In practice, VAEs usually require multiple training rounds to choose the amount of information the latent…

Machine Learning · Computer Science 2023-08-21 Juhan Bae , Michael R. Zhang , Michael Ruan , Eric Wang , So Hasegawa , Jimmy Ba , Roger Grosse

We present new intuitions and theoretical assessments of the emergence of disentangled representation in variational autoencoders. Taking a rate-distortion theory perspective, we show the circumstances under which representations aligned…

Linear attention reduces the quadratic cost of softmax attention to $\mathcal{O}(T)$, but its memory state grows as $\mathcal{O}(T)$ in Frobenius norm, causing progressive interference between stored associations. We introduce…

Machine Learning · Computer Science 2026-05-13 Vishal Pandey , Gopal Singh
‹ Prev 1 2 3 10 Next ›