Related papers: Learning the Arrow of Time
Uncovering the origin of the arrow of time remains a fundamental scientific challenge. Within the framework of statistical physics, this problem was inextricably associated with the second law of thermodynamics, which declares that entropy…
Temporal resolution of visual information processing is thought to be an important factor in predator-prey interactions, shaped in the course of evolution by animals' ecology. Here I show that light can be considered to have a dual role of…
This paper presents an online method that learns optimal decisions for a discrete time Markov decision problem with an opportunistic structure. The state at time $t$ is a pair $(S(t),W(t))$ where $S(t)$ takes values in a finite set…
We clarify and strengthen our demonstration that arrows of time necessarily arise in unconfined systems. Contrary to a recent claim, this does not require an improbable selection principle.
In this paper, we present an online reinforcement learning algorithm for constrained Markov decision processes with a safety constraint. Despite the necessary attention of the scientific community, considering stochastic stopping time, the…
Active learning agents typically employ a query selection algorithm which solely considers the agent's learning objectives. However, this may be insufficient in more realistic human domains. This work uses imitation learning to enable an…
The purpose of this paper is to highlight the central role that the time asymmetry of stability plays in feedback control. We show that this provides a new perspective on the use of doubly-infinite or semi-infinite time axes for signal…
According to the dominant view, time in perceptual decision making is used for integrating new sensory evidence. Based on a probabilistic framework, we investigated the alternative hypothesis that time is used for gradually refining an…
The ability to predict the future in a given domain can be acquired by discovering empirically from experience certain temporal patterns that tend to repeat unerringly. Previous works in time series analysis allow one to make quantitative…
The arrow of time dilemma: the laws of physics are invariant for time inversion, whereas the familiar phenomena we see everyday are not (i.e. entropy increases). I show that, within a quantum mechanical framework, all phenomena which leave…
Constructing an accurate system model for formal model verification can be both resource demanding and time-consuming. To alleviate this shortcoming, algorithms have been proposed for automatically learning system models based on observed…
Reinforcement learning (RL) allows to solve complex tasks such as Go often with a stronger performance than humans. However, the learned behaviors are usually fixed to specific tasks and unable to adapt to different contexts. Here we…
This essay offers a meta-level analysis in the sociology and history of physics in the context of the so-called "Arrow of Time Problem" or "Two Times Problem," which asserts that the empirically observed directionality of time is in…
What is the physical origin of the arrow of time? It is a commonly held belief in the physics community that it relates to the increase of entropy as it appears in the statistical interpretation of the second law of thermodynamics. At the…
We investigate the statistical arrow of time for a quantum system being monitored by a sequence of measurements. For a continuous qubit measurement example, we demonstrate that time-reversed evolution is always physically possible, provided…
A model quantum cosmology is used to illustrate how arrows of time emerge in a universe governed by a time-neutral dynamical theory constrained by time asymmetric initial and final boundary conditions represented by initial and final…
We address the problem of reinforcement learning in which observations may exhibit an arbitrary form of stochastic dependence on past observations and actions. The task for an agent is to attain the best possible asymptotic reward where the…
We derive and analyze learning algorithms for apprenticeship learning, policy evaluation, and policy gradient for average reward criteria. Existing algorithms explicitly require an upper bound on the mixing time. In contrast, we build on…
We consider the problem of how strategic users with asymmetric information can learn an underlying time varying state in a user-recommendation system. Users who observe private signals about the state, sequentially make a decision about…
This work explores the implications of assuming time symmetry and applying bridge-type, time-symmetric temporal boundary conditions to deterministic laws of nature with random components. The analysis, drawing on the works of Kolmogorov and…