Related papers: Why Should I Trust You, Bellman? The Bellman Error…
We prove performance guarantees of two algorithms for approximating $Q^\star$ in batch reinforcement learning. Compared to classical iterative methods such as Fitted Q-Iteration---whose performance loss incurs quadratic dependence on…
Predictive mean matching (PMM) is a popular imputation strategy that imputes missing values by borrowing observed values from other cases with similar expectations. We show that, unlike other imputation strategies, PMM is not guaranteed to…
In this paper we formulate and study an optimal switching problem under partial information. In our model the agent/manager/investor attempts to maximize the expected reward by switching between different states/investments. However, he is…
This paper develops an inverse reinforcement learning algorithm aimed at recovering a reward function from the observed actions of an agent. We introduce a strategy to flexibly handle different types of actions with two approximations of…
Bell's theorem is typically understood as the proof that quantum theory is incompatible with local-hidden-variable models. More generally, we can see the violation of a Bell inequality as witnessing the impossibility of explaining quantum…
We consider the differentiation of the value function for parametric optimization problems. Such problems are ubiquitous in Machine Learning applications such as structured support vector machines, matrix factorization and min-min or…
In this paper we are concerned with the error-covariance lower-bounding problem in Kalman filtering: a sensor releases a set of measurements to the data fusion/estimation center, which has a perfect knowledge of the dynamic model, to allow…
This paper establishes a rigorous connection between regularized discrete-time reinforcement learning (RL) and continuous-time stochastic optimal control. Specifically, classical RL algorithms are typically solving a regularized…
The statistics behind Bell's inequality is demonstrated to allow a Kolmogorovian (i.e. classical) model of probabilities that recovers the quantum covariance.
Calibration$\unicode{x2014}$the problem of ensuring that predicted probabilities align with observed class frequencies$\unicode{x2014}$is a basic desideratum for reliable prediction with machine learning systems. Calibration error is…
In a recent simulation study, Goodman et al. (2019) compare several methods with regard to their type I and type II error rates in case of a thick null hypothesis that includes all values that are practically equivalent to the point null…
We study finite-horizon continuous-time policy evaluation from discrete closed-loop trajectories under time-inhomogeneous dynamics. The target value surface solves a backward parabolic equation, but the Bellman baseline obtained from…
Mermin states that his nontechnical version of Bell's theorem stands and is not invalidated by time and setting dependent instrument parameters as claimed in one of our previous papers. We identify deviations from well-established protocol…
Batch reinforcement learning (RL) is important to apply RL algorithms to many high stakes tasks. Doing batch RL in a way that yields a reliable new policy in large domains is challenging: a new decision policy may visit states and actions…
In this paper, we consider discrete-time infinite horizon problems of optimal control to a terminal set of states. These are the problems that are often taken as the starting point for adaptive dynamic programming. Under very general…
It is widely accepted that the violation of Bell inequalities excludes local theories of the quantum realm. This paper presents a new derivation of the inequalities from non-trivial non-local theories and formulates a stronger Bell argument…
We formally prove the existence of an enduring incongruence pervading a widespread interpretation of the Bell inequality and explain how to rationally avoid it with a natural assumption justified by explicit reference to a mathematical…
Reinforcement learning (RL) algorithms assume that users specify tasks by manually writing down a reward function. However, this process can be laborious and demands considerable technical expertise. Can we devise RL algorithms that instead…
This paper presents a theory of error in cross-validation testing of algorithms for predicting real-valued attributes. The theory justifies the claim that predicting real-valued attributes requires balancing the conflicting demands of…
Measures of accuracy usually score how accurate a specified credence depending on whether the proposition is true or false. A key requirement for such measures is strict propriety; that probabilities expect themselves to be most accurate.…