相关论文: Reconciling Early and Late Time Tensions with Rein…
The Hubble tension seems to be a crisis with $\sim5\sigma$ discrepancy between the most recent local distance ladder measurement from type Ia supernovae calibrated by Cepheids and the global fitting constraint from the cosmic microwave…
An only early or only late time alteration to $\Lambda$CDM has been inadequate at resolving both the $H_0$ and $S_8$ tensions simultaneously; however, a combination of early and late time alterations to $\Lambda$CDM can provide a solution…
The considerable difference between early and late universe measurements of the Hubble constant, called the Hubble tension, poses a potential challenge to the standard $\Lambda$CDM cosmological model. We examine an interacting dark…
Reinforcement learning is one of the core components in designing an artificial intelligent system emphasizing real-time response. Reinforcement learning influences the system to take actions within an arbitrary environment either having…
Many late time approaches for the solution of the Hubble tension use late time smooth deformations of the Hubble expansion rate $H(z)$ of the Planck18/$\Lambda$CDM best fit to match the locally measured value of $H_0$ while effectively…
Temporal difference (TD) methods constitute a class of methods for learning predictions in multi-step prediction problems, parameterized by a recency factor lambda. Currently the most important application of these methods is to temporal…
Recently, it has been proposed that Hubble tension can be addressed in the $\Lambda$CDM model if the lookback time approach is considered on the redshift $z$ measured. From this interesting proposal, the lookback time evolution seems to…
Machine learning models are often used at test-time subject to constraints and trade-offs not present at training-time. For example, a computer vision model operating on an embedded device may need to perform real-time inference, or a…
Deep reinforcement learning enables algorithms to learn complex behavior, deal with continuous action spaces and find good strategies in environments with high dimensional state spaces. With deep reinforcement learning being an active area…
Commonly in reinforcement learning (RL), rewards are discounted over time using an exponential function to model time preference, thereby bounding the expected long-term reward. In contrast, in economics and psychology, it has been shown…
There has been a significant interest in modifications of the standard $\Lambda$ Cold Dark Matter ($\Lambda$CDM) cosmological model prompted by tensions between certain datasets, most notably the Hubble tension. The late-time modifications…
Accuracy and timeliness are indeed often conflicting goals in prediction tasks. Premature predictions may yield a higher rate of false alarms, whereas delaying predictions to gather more information can render them too late to be useful. In…
The standard cosmological model successfully describes many observations from widely different epochs of the Universe, from primordial nucleosynthesis to the accelerating expansion of the present day. However, as the basic cosmological…
Reinforcement learning (RL) is a promising approach for aligning large language models (LLMs) knowledge with sequential decision-making tasks. However, few studies have thoroughly investigated the impact on LLM agents capabilities of…
Many problems in astrophysics cover multiple orders of magnitude in spatial and temporal scales. While simulating systems that experience rapid changes in these conditions, it is essential to adapt the (time-) step size to capture the…
There are two distinct approaches to solving reinforcement learning problems, namely, searching in value function space and searching in policy space. Temporal difference methods and evolutionary algorithms are well-known examples of these…
We construct data-driven solutions to the Hubble tension which are perturbative modifications to the fiducial $\Lambda$CDM cosmology, using the Fisher bias formalism. Taking as proof of principle the case of a time-varying electron mass and…
We present one of the first algorithms on model based reinforcement learning and trajectory optimization with free final time horizon. Grounded on the optimal control theory and Dynamic Programming, we derive a set of backward differential…
This paper proposes a new reinforcement learning with hyperbolic discounting. Combining a new temporal difference error with the hyperbolic discounting in recursive manner and reward-punishment framework, a new scheme to learn the optimal…
We present an end-to-end framework for the Assignment Problem with multiple tasks mapped to a group of workers, using reinforcement learning while preserving many constraints. Tasks and workers have time constraints and there is a cost…