Related papers: Exercising Control When Confronted by a (Brownian)…
I present arguments against the hypothesis put forward by Silver, Singh, Precup, and Sutton ( https://www.sciencedirect.com/science/article/pii/S0004370221000862 ) : reward maximization is not enough to explain many activities associated…
We introduce a dynamic model in which a developer incrementally improves a product of uncertain quality over time, with the quality evolving as a controlled Brownian motion. At each moment in time, the developer can continue exploring by…
In the field of reinforcement learning there has been recent progress towards safety and high-confidence bounds on policy performance. However, to our knowledge, no practical methods exist for determining high-confidence policy performance…
The recent paper `"Reward is Enough" by Silver, Singh, Precup and Sutton posits that the concept of reward maximisation is sufficient to underpin all intelligence, both natural and artificial. We contest the underlying assumption of Silver…
Contextual dueling bandit is used to model the bandit problems, where a learner's goal is to find the best arm for a given context using observed noisy human preference feedback over the selected arms for the past contexts. However,…
A simple random walk and a Brownian motion are considered on a spider that is a collection of half lines (we call them legs) joined in the origin. We give a strong approximation of these two objects and their local times. For fixed number…
We propose a general framework for studying optimal impulse control problem in the presence of uncertainty on the parameters. Given a prior on the distribution of the unknown parameters, we explain how it should evolve according to the…
We investigate a class of optimal stopping problems arising in, for example, studies considering the timing of an irreversible investment when the underlying follows a skew Brownian motion. Our results indicate that the local directional…
This paper introduces dynamic mechanism design in an elementary fashion. We first examine optimal dynamic mechanisms: We find necessary and sufficient conditions for perfect Bayesian incentive compatibility and formulate the optimal dynamic…
We address the problem of identifying the optimal policy with a fixed confidence level in a multi-armed bandit setup, when \emph{the arms are subject to linear constraints}. Unlike the standard best-arm identification problem which is well…
We study the experimentation dynamics of a decision maker (DM) in a two-armed bandit setup (Bolton and Harris (1999)), where the agent holds ambiguous beliefs regarding the distribution of the return process of one arm and is certain about…
Sequential decision making under uncertainty is studied in a mixed observability domain. The goal is to maximize the amount of information obtained on a partially observable stochastic process under constraints imposed by a fully observable…
This paper presents a model of pedestrian crossing decisions, based on the theory of computational rationality. It is assumed that crossing decisions are boundedly optimal, with bounds on optimality arising from human cognitive limitations.…
We integrate dual-process theories of human cognition with evolutionary game theory to study the evolution of automatic and controlled decision-making processes. We introduce a model where agents who make decisions using either automatic or…
Dynamical systems are frequently used to model biological systems. When these models are fit to data it is necessary to ascertain the uncertainty in the model fit. Here we present prediction deviation, a new metric of uncertainty that…
We present a new algorithm for computing upper bounds on the number of executions of each program instruction during any single program run. The upper bounds are expressed as functions of program input values. The algorithm is primarily…
In this paper, we study a class of stochastic optimal control problem with jumps under partial information. More precisely, the controlled systems are described by a fully coupled nonlinear multi- dimensional forward-backward stochastic…
We consider the Brownian ``spider process'', also known as Walsh Brownian motion, first introduced in the epilogue of Walsh 1978. The paper provides the best constant $C_n$ for the inequality $$ E D_\tau\leq C_n \sqrt{E \tau},$$ where…
Individual decision-makers consume information revealed by the previous decision makers, and produce information that may help in future decisions. This phenomenon is common in a wide range of scenarios in the Internet economy, as well as…
We study a signaling game between an employer and a potential employee, where the employee has private information regarding their production capacity. At the initial stage, the employee communicates a salary claim, after which the true…