Related papers: Rational hyperbolic discounting
A set of objects is to be divided fairly among agents with different tastes, modeled by additive utility-functions. If we consider the objects as indivisible, many instances of the decision problem: ``Is there a fair division of the objects…
In reinforcement learning (RL), the goal is to obtain an optimal policy, for which the optimality criterion is fundamentally important. Two major optimality criteria are average and discounted rewards. While the latter is more popular, it…
Our goal is for agents to optimize the right reward function, despite how difficult it is for us to specify what that is. Inverse Reinforcement Learning (IRL) enables us to infer reward functions from demonstrations, but it usually assumes…
Fair division has long been an important problem in the economics literature. In this note, we consider the existence of proportionally fair allocations of indivisible goods, i.e., allocations of indivisible goods in which every agent gets…
This paper considers an infinite-horizon Markov decision process (MDP) that allows for general non-exponential discount functions, in both discrete and continuous time. Due to the inherent time inconsistency, we look for a randomized…
Inflation is painful, for firms, customers, employees, and society. But careful study of periods of hyperinflation point to ways that firms can adapt. In particular, companies need to think about how to change prices regularly and cheaply,…
An ability to postpone one's execution without penalty provides an important strategic advantage in high-frequency trading. To elucidate competition between traders one has to formulate to a quantitative theory of formation of the execution…
To determine the welfare implications of price changes in demand data, we introduce a revealed preference relation over prices. We show that the absence of cycles in this relation characterizes a consumer who trades off the utility of…
We review the theory of renewal reward processes, which describes renewal processes that have some cost or reward associated with each cycle. We present a new simplified proof of the renewal reward theorem that mimics the proof of the…
Consider a game where Alice generates an integer and Bob wins if he can factor that integer. Traditional game theory tells us that Bob will always win this game even though in practice Alice will win given our usual assumptions about the…
One method to offer some bidders a discount in a first-price auction is to augment their bids when selecting a winner but only charge them their original bids should they win. Another method is to use their original bids to select a winner,…
I provide a model of rational inattention with heterogeneity and prove it is observationally equivalent to a state-dependent stochastic choice model subject to attention costs. I demonstrate that additive separability of unobservable…
Incentives have surprisingly inconsistent effects when it comes to encouraging people to behave prosocially. Classical economic theory, according to which a specific behavior becomes more prevalent when it is rewarded, struggles to explain…
In repeated games, cooperation is possible in equilibrium only if players are sufficiently patient, and long-term gains from cooperation outweigh short-term gains from deviation. What happens if the players have incomplete information…
The consumption function maps current wealth and the exogenous state to current consumption. We prove the existence and uniqueness of a consumption function when the agent has a preference for wealth. When the period utility functions are…
Generally accepted depreciation methods do not compute the intrinsic value of an asset, as they do not factor for the Time Value of Money, a key principle within financial theory. This is disadvantageous, as knowing the intrinsic value of…
Problem definition: Mining for heterogeneous responses to an intervention is a crucial step for data-driven operations, for instance to personalize treatment or pricing. We investigate how to estimate price sensitivity from…
LLMs utilizing chain-of-thought reasoning often waste substantial compute by producing long, incorrect responses. Abstention can mitigate this by withholding outputs unlikely to be correct. While most abstention methods decide to withhold…
The absent-minded driver's problem illustrates that probabilistic strategies can give higher pay-offs than deterministic ones. We show that there are strategies using quantum entangled states that give even higher pay-offs, both for the…
We present a model of optimal training of a rational, sluggish agent. A trainer commits to a discrete-time, finite-state Markov process that governs the evolution of training intensity. Subsequently, the agent monitors the state and adjusts…