English
Related papers

Related papers: When perceptual time stands still: Long stable mem…

200 papers

Linear TD($\lambda$) is one of the most fundamental reinforcement learning algorithms for policy evaluation. Previously, convergence rates are typically established under the assumption of linearly independent features, which does not hold…

Machine Learning · Computer Science 2025-10-15 Zixuan Xie , Xinyu Liu , Rohan Chandra , Shangtong Zhang

Off-policy learning enables a reinforcement learning (RL) agent to reason counterfactually about policies that are not executed and is one of the most important ideas in RL. It, however, can lead to instability when combined with function…

Machine Learning · Computer Science 2025-03-03 Xiaochi Qian , Shangtong Zhang

The recency heuristic in reinforcement learning is the assumption that stimuli that occurred closer in time to an acquired reward should be more heavily reinforced. The recency heuristic is one of the key assumptions made by TD($\lambda$),…

Machine Learning · Computer Science 2024-08-27 Brett Daley , Marlos C. Machado , Martha White

Imitating successful behavior is a natural and frequently applied approach to trust in when facing scenarios for which we have little or no experience upon which we can base our decision. In this paper, we consider such behavior in atomic…

Computer Science and Game Theory · Computer Science 2008-10-04 Heiner Ackermann , Petra Berenbrink , Simon Fischer , Martin Hoefer

We establish new results for estimation and inference in financial durations models, where events are observed over a given time span, such as a trading day, or a week. For the classical autoregressive conditional duration (ACD) models by…

Econometrics · Economics 2022-12-02 Giuseppe Cavaliere , Thomas Mikosch , Anders Rahbek , Frederik Vilandt

The Fundamental Theorem of Language Change (Yang, 2000) implies the impossibility of stable variation in the Variational Learning framework, but only in the special case where two, and not more, grammatical variants compete. Introducing the…

Artificial Intelligence · Computer Science 2020-03-16 Henri Kauhanen

We show that trial-to-trial variability in sensory detection of a weak visual stimulus is dramatically diminished when rather than presenting a fixed stimulus contrast, fluctuations in a subject's judgment are matched by fluctuations in…

Neurons and Cognition · Quantitative Biology 2011-03-29 Shimon Marom , Avner Wallach

In the optional prisoner's dilemma (OPD), players can choose to cooperate and defect as usual, but can also abstain as a third possible strategy. This strategy models the players' participation in the game and is a relevant aspect in many…

Computer Science and Game Theory · Computer Science 2021-06-15 Leonardo Stella , Dario Bauso

With deep learning approaches becoming state-of-the-art in many speech (as well as non-speech) related machine learning tasks, efforts are being taken to delve into the neural networks which are often considered as a black box. In this…

Machine Learning · Computer Science 2018-08-27 Jeroen Zegers , Hugo Van hamme

Reinforcement learning lies at the intersection of several challenges. Many applications of interest involve extremely large state spaces, requiring function approximation to enable tractable computation. In addition, the learner has only a…

Machine Learning · Computer Science 2021-05-11 Andrew Jacobsen , Alan Chan

We study computational and statistical aspects of learning Latent Markov Decision Processes (LMDPs). In this model, the learner interacts with an MDP drawn at the beginning of each epoch from an unknown mixture of MDPs. To sidestep known…

Machine Learning · Computer Science 2024-06-13 Fan Chen , Constantinos Daskalakis , Noah Golowich , Alexander Rakhlin

Temporal-difference (TD) learning is widely regarded as one of the most popular algorithms in reinforcement learning (RL). Despite its widespread use, it has only been recently that researchers have begun to actively study its finite time…

Machine Learning · Computer Science 2025-04-16 Han-Dong Lim , Donghwan Lee

Reviewing plays an important role when learning knowledge. The knowledge acquisition at a certain time point may be strongly inspired with the help of previous experience. Thus the knowledge growing procedure should show strong relationship…

Computer Vision and Pattern Recognition · Computer Science 2025-05-13 Dongwei Wang , Zhi Han , Yanmei Wang , Xiai Chen , Baichen Liu , Yandong Tang

Two nested classes of discrete-time linear time-invariant systems, which differ by the set of periodic signals that they leave invariant, are studied. The first class preserves the property of periodic monotonicity (period-wise…

Optimization and Control · Mathematics 2026-02-10 Christian Grussler

We present a theoretical study aiming at model fitting for sensory neurons. Conventional neural network training approaches are not applicable to this problem due to lack of continuous data. Although the stimulus can be considered as a…

Neurons and Cognition · Quantitative Biology 2017-09-28 R. Ozgur Doruk , Kechen Zhang

Variability in neural responses is an ubiquitous phenomenon in neurons, usually modeled with stochastic differential equations. In particular, stochastic integrate-and-fire models are widely used to simplify theoretical studies. The…

Neurons and Cognition · Quantitative Biology 2009-06-12 Eugenio Urdapilleta , Ines Samengo

A tenet of reinforcement learning is that the agent always observes rewards. However, this is not true in many realistic settings, e.g., a human observer may not always be available to provide rewards, sensors may be limited or…

Machine Learning · Computer Science 2026-03-24 Alireza Kazemipour , Simone Parisi , Matthew E. Taylor , Michael Bowling

Temporal difference learning and Residual Gradient methods are the most widely used temporal difference based learning algorithms; however, it has been shown that none of their objective functions is optimal w.r.t approximating the true…

Machine Learning · Computer Science 2017-04-21 Bo Liu , Daoming Lyu , Wen Dong , Saad Biaz

We study two-player zero-sum stochastic games, and propose a form of independent learning dynamics called Doubly Smoothed Best-Response dynamics, which integrates a discrete and doubly smoothed variant of the best-response dynamics into…

Computer Science and Game Theory · Computer Science 2023-03-07 Zaiwei Chen , Kaiqing Zhang , Eric Mazumdar , Asuman Ozdaglar , Adam Wierman

This article is dedicated to the study and comparison of two chemostat-like competition models in a heterogeneous environment. The first model is a probabilistic model where we build a PDMP simulating the effect of the temporal…

Dynamical Systems · Mathematics 2018-06-29 Sten Madec , G Lagasquie
‹ Prev 1 4 5 6 7 8 10 Next ›