English
Related papers

Related papers: Exercising Control When Confronted by a (Brownian)…

200 papers

I present arguments against the hypothesis put forward by Silver, Singh, Precup, and Sutton ( https://www.sciencedirect.com/science/article/pii/S0004370221000862 ) : reward maximization is not enough to explain many activities associated…

Artificial Intelligence · Computer Science 2024-11-12 Vacslav Glukhov

We introduce a dynamic model in which a developer incrementally improves a product of uncertain quality over time, with the quality evolving as a controlled Brownian motion. At each moment in time, the developer can continue exploring by…

Theoretical Economics · Economics 2025-07-22 Santiago Oliveros

In the field of reinforcement learning there has been recent progress towards safety and high-confidence bounds on policy performance. However, to our knowledge, no practical methods exist for determining high-confidence policy performance…

Artificial Intelligence · Computer Science 2018-06-26 Daniel S. Brown , Scott Niekum

The recent paper `"Reward is Enough" by Silver, Singh, Precup and Sutton posits that the concept of reward maximisation is sufficient to underpin all intelligence, both natural and artificial. We contest the underlying assumption of Silver…

Contextual dueling bandit is used to model the bandit problems, where a learner's goal is to find the best arm for a given context using observed noisy human preference feedback over the selected arms for the past contexts. However,…

Machine Learning · Computer Science 2025-04-17 Arun Verma , Zhongxiang Dai , Xiaoqiang Lin , Patrick Jaillet , Bryan Kian Hsiang Low

A simple random walk and a Brownian motion are considered on a spider that is a collection of half lines (we call them legs) joined in the origin. We give a strong approximation of these two objects and their local times. For fixed number…

Probability · Mathematics 2017-05-12 Endre Csaki , Miklos Csorgo , Antonia Foldes , Pal Revesz

We propose a general framework for studying optimal impulse control problem in the presence of uncertainty on the parameters. Given a prior on the distribution of the unknown parameters, we explain how it should evolve according to the…

Probability · Mathematics 2017-12-06 N. Baradel , B. Bouchard , Ngoc Minh Dang

We investigate a class of optimal stopping problems arising in, for example, studies considering the timing of an irreversible investment when the underlying follows a skew Brownian motion. Our results indicate that the local directional…

Probability · Mathematics 2016-08-17 Luis H. R. Alvarez E. , Paavo Salminen

This paper introduces dynamic mechanism design in an elementary fashion. We first examine optimal dynamic mechanisms: We find necessary and sufficient conditions for perfect Bayesian incentive compatibility and formulate the optimal dynamic…

Theoretical Economics · Economics 2024-09-10 Kiho Yoon

We address the problem of identifying the optimal policy with a fixed confidence level in a multi-armed bandit setup, when \emph{the arms are subject to linear constraints}. Unlike the standard best-arm identification problem which is well…

Machine Learning · Computer Science 2024-01-26 Emil Carlsson , Debabrota Basu , Fredrik D. Johansson , Devdatt Dubhashi

We study the experimentation dynamics of a decision maker (DM) in a two-armed bandit setup (Bolton and Harris (1999)), where the agent holds ambiguous beliefs regarding the distribution of the return process of one arm and is certain about…

Theoretical Economics · Economics 2021-04-02 Farzad Pourbabaee

Sequential decision making under uncertainty is studied in a mixed observability domain. The goal is to maximize the amount of information obtained on a partially observable stochastic process under constraints imposed by a fully observable…

Artificial Intelligence · Computer Science 2016-03-16 Mikko Lauri , Risto Ritala

This paper presents a model of pedestrian crossing decisions, based on the theory of computational rationality. It is assumed that crossing decisions are boundedly optimal, with bounds on optimality arising from human cognitive limitations.…

Artificial Intelligence · Computer Science 2024-02-08 Yueyang Wang , Aravinda Ramakrishnan Srinivasan , Jussi P. P. Jokinen , Antti Oulasvirta , Gustav Markkula

We integrate dual-process theories of human cognition with evolutionary game theory to study the evolution of automatic and controlled decision-making processes. We introduce a model where agents who make decisions using either automatic or…

Dynamical Systems · Mathematics 2015-07-07 Danielle F. P. Toupo , Steven H. Strogatz , Jonathan D. Cohen , David G. Rand

Dynamical systems are frequently used to model biological systems. When these models are fit to data it is necessary to ascertain the uncertainty in the model fit. Here we present prediction deviation, a new metric of uncertainty that…

Applications · Statistics 2017-06-08 Benjamin Letham , Portia A. Letham , Cynthia Rudin , Edward P. Browne

We present a new algorithm for computing upper bounds on the number of executions of each program instruction during any single program run. The upper bounds are expressed as functions of program input values. The algorithm is primarily…

Programming Languages · Computer Science 2016-05-13 Pavel Čadek , Jan Strejček , Marek Trtík

In this paper, we study a class of stochastic optimal control problem with jumps under partial information. More precisely, the controlled systems are described by a fully coupled nonlinear multi- dimensional forward-backward stochastic…

Optimization and Control · Mathematics 2009-11-18 Qingxin Meng

We consider the Brownian ``spider process'', also known as Walsh Brownian motion, first introduced in the epilogue of Walsh 1978. The paper provides the best constant $C_n$ for the inequality $$ E D_\tau\leq C_n \sqrt{E \tau},$$ where…

Probability · Mathematics 2021-06-14 Ewelina Bednarz , Philip A. Ernst , Adam Osekowski

Individual decision-makers consume information revealed by the previous decision makers, and produce information that may help in future decisions. This phenomenon is common in a wide range of scenarios in the Internet economy, as well as…

Computer Science and Game Theory · Computer Science 2019-05-06 Yishay Mansour , Aleksandrs Slivkins , Vasilis Syrgkanis

We study a signaling game between an employer and a potential employee, where the employee has private information regarding their production capacity. At the initial stage, the employee communicates a salary claim, after which the true…

Optimization and Control · Mathematics 2024-06-26 Erik Ekström , Topias Tolonen-Weckström