Related papers: Why Should I Trust You, Bellman? The Bellman Error…
With the advent of AI technologies, humans and robots are increasingly teaming up to perform collaborative tasks. To enable smooth and effective collaboration, the topic of value alignment (operationalized herein as the degree of dynamic…
In real data, missing values occur frequently, which affects the interpretation with interpretable machine learning (IML) methods. Recent work considers bias and shows that model explanations may differ between imputation methods, while…
We investigate a classical statistical model and show that Mermin's version of a Bell inequality is violated. We get this violation, if the measurement modifies the ensemble, a feature, which is also characteristic for measurement processes…
We precisely compute the Bellman function of two variables of the dyadic maximal operator in relation to Kolmogorov inequality. In this way we give an alternative proof of the results in [5].Additionally, we characterize the sequences of…
Deep metric learning techniques have been used for visual representation in various supervised and unsupervised learning tasks through learning embeddings of samples with deep networks. However, classic approaches, which employ a fixed…
Offline reinforcement learning promises policy improvement from logged interaction data alone, yet state-of-the-art algorithms remain vulnerable to value over-estimation and to violations of domain knowledge such as monotonicity or…
The Bregman divergence (Bregman distance, Bregman measure of distance) is a certain useful substitute for a distance, obtained from a well-chosen function (the "Bregman function"). Bregman functions and divergences have been extensively…
Markov decision problems are most commonly solved via dynamic programming. Another approach is Bellman residual minimization, which directly minimizes the squared Bellman residual objective function. However, compared to dynamic…
In this paper, we provide an example of the optimal growth model in which there exist infinitely many solutions to the Hamilton-Jacobi-Bellman equation but the value function does not satisfy this equation. We consider the cause of this…
There are several versions of Bell's inequalities, proved in different contexts, using different sets of assumptions. The discussions of their experimental violation often disregard some required assumptions and use loose formulations of…
We study the Bellman equation in the Wasserstein space arising in the study of mean field control problems, namely stochastic optimal control problems for McKean-Vlasov diffusion processes.Using the standard notion of viscosity solution \`a…
We describe a nonlinear generalization of dual dynamic programming theory and its application to value function estimation for deterministic control problems over continuous state and action spaces, in a discrete-time infinite horizon…
We present a bound for value-prediction error with respect to model misspecification that is tight, including constant factors. This is a direct improvement of the "simulation lemma," a foundational result in reinforcement learning. We…
Traditional statements of the celebrated Kalman filter algorithm focus on the estimation of state, but not the output. For any outputs, measured or auxiliary, it is usually assumed that the posterior state estimates and known inputs are…
Subadditive set functions play a pivotal role in computational economics (especially in combinatorial auctions), combinatorial optimization or artificial intelligence applications such as interpretable machine learning. However, specifying…
Suppose we know that an object is in a sorted table and we want to determine the index of that object. To achieve this goal we could perform a binary search. However, suppose it is time-consuming to determine the relative position of that…
Gradient descent or its variants are popular in training neural networks. However, in deep Q-learning with neural network approximation, a type of reinforcement learning, gradient descent (also known as Residual Gradient (RG)) is barely…
This paper explores the application of nonsmooth analysis in the Wasserstein space to finite-horizon optimal control problems for nonlocal continuity equations. We characterize the value function as a strict viscosity solution of the…
Finite difference approximations to multi-asset American put option price are considered. The assets are modelled as a multi-dimensional diffusion process with variable drift and volatility. Approximation error of order one quarter with…
Despite its experimental success, Model-based Reinforcement Learning still lacks a complete theoretical understanding. To this end, we analyze the error in the cumulative reward using a contraction approach. We consider both stochastic and…