Related papers: Risk and Intertemporal Preferences over Time Lotte…
In this letter, the optimality of the likelihood ratio test (LRT) is investigated for binary hypothesis testing problems in the presence of a behavioral decision-maker. By utilizing prospect theory, a behavioral decision-maker is modeled to…
We present a framework to interpret signal temporal logic (STL) formulas over discrete-time stochastic processes in terms of the induced risk. Each realization of a stochastic process either satisfies or violates an STL formula. In fact, we…
RL is increasingly being used to control robotic systems that interact closely with humans. This interaction raises the problem of safe RL: how to ensure that a RL-controlled robotic system never, for instance, injures a human. This problem…
Ensuring that AI systems make strategic decisions aligned with the specified preferences in adversarial sequential interactions is a critical challenge for developing trustworthy AI systems, especially when the environment is stochastic and…
In this paper, we consider a class of stochastic optimal control problems with risk constraints that are expressed as bounded probabilities of failure for particular initial states. We present here a martingale approach that diffuses a risk…
This paper axiomatizes, in a two-stage setup, a new theory for decision under risk and ambiguity. The axiomatized preference relation $\succeq$ on the space $\tilde{V}$ of random variables induces an ambiguity index $c$ on the space…
For AI systems to be useful to humans, they must understand and act in accordance with our values and preferences. Since specifying preferences is a hard task, inverse reinforcement learning (IRL) aims to develop methods that allow for…
In the ultimatum game, the challenge is to explain why responders reject non-zero offers thereby defying classical rationality. Fairness and related notions have been the main explanations so far. We explain this rejection behavior via the…
The theory of rational choice assumes that when people make decisions they do so in order to maximize their utility. In order to achieve this goal they ought to use all the information available and consider all the choices available to…
This paper investigates the use of risk measures and theories of choice for modeling risk-averse route choice and traffic network equilibrium with random travel times. We interpret the postulates of these theories in the context of routing,…
Reinforcement learning (RL) is central to improving reasoning in large language models (LLMs) but typically requires ground-truth rewards. Test-Time Reinforcement Learning (TTRL) removes this need by using majority-vote rewards, but relies…
The role of specific cognitive processes in deviations from constant discounting in intertemporal choice is not well understood. We evaluated decreased impatience in intertemporal choice tasks independent of discounting rate and…
The use of hypothetical instead of real decision-making incentives remains under debate after decades of economic experiments. Standard incentivized experiments involve substantial monetary costs due to participants' earnings and often…
The equity risk premium puzzle is that the return on equities has far exceeded the average return on short-term risk-free debt and cannot be explained by conventional representative-agent consumption based equilibrium models. We review a…
Inverse reinforcement learning (IRL) aims to recover the reward function and the associated optimal policy that best fits observed sequences of states and actions implemented by an expert. Many algorithms for IRL have an inherently nested…
We investigate the performance of the Kelly rule in a setting in which the dynamics of the return is represented by a time change process. We find that in this general semi-martingale setting the Kelly rule does not maximize the average…
Bellman formulated a vague principle for optimization over time, which characterizes optimal policies by stating that a decision maker should not regret previous decisions retrospectively. This paper addresses time consistency in stochastic…
One typical assumption in inverse reinforcement learning (IRL) is that human experts act to optimize the expected utility of a stochastic cost with a fixed distribution. This assumption deviates from actual human behaviors under ambiguity.…
Distributional reinforcement learning (RL) is a powerful framework increasingly adopted in safety-critical domains for its ability to optimize risk-sensitive objectives. However, the role of the discount factor is often overlooked, as it is…
Timing decisions are common: when to file your taxes, finish a referee report, or complete a task at work. We ask whether time preferences can be inferred when \textsl{only} task completion is observed. To answer this question, we analyze…