Related papers: Why Should I Trust You, Bellman? The Bellman Error…
Given that AI systems are set to play a pivotal role in future decision-making processes, their trustworthiness and reliability are of critical concern. Due to their scale and complexity, modern AI systems resist direct interpretation, and…
We study optimality for the safety-constrained Markov decision process which is the underlying framework for safe reinforcement learning. Specifically, we consider a constrained Markov decision process (with finite states and finite…
This short article concentrates on the conceptual aspects of the violation of Bell inequalities, and acts as a map to the 265 cited references. The article outlines (a) relevant characteristics of quantum mechanics, such as statistical…
Missing values are a fundamental problem in data science. Many datasets have missing values that must be properly handled because the way missing values are treated can have large impact on the resulting machine learning model. In medical…
Bell inequalities were meant to test quantum mechanics vs local hidden variable models, but can also be used to verify entanglement. For entanglement verification purposes one assumes the validity of quantum mechanics as well as quantum…
Recent work on policy learning from observational data has highlighted the importance of efficient policy evaluation and has proposed reductions to weighted (cost-sensitive) classification. But, efficient policy evaluation need not yield…
In reinforcement learning the Q-values summarize the expected future rewards that the agent will attain. However, they cannot capture the epistemic uncertainty about those rewards. In this work we derive a new Bellman operator with…
The interpretation of the meaning of Quantum Mechanics has faced controversy since its inception. Bell's inequalities are a touchstone in this controversy. Their observed violation demonstrates that at least one of the hypotheses involved…
We study offline reinforcement learning (RL) which seeks to learn a good policy based on a fixed, pre-collected dataset. A fundamental challenge behind this task is the distributional shift due to the dataset lacking sufficient exploration,…
Many safety failures in machine learning arise when models are used to assign predictions to people (often in settings like lending, hiring, or content moderation) without accounting for how individuals can change their inputs. In this…
The error autocorrection effect means that in a calculation all the intermediate errors compensate each other, so the final result is much more accurate than the intermediate results. In this case standard interval estimates are too…
We study online transfer reinforcement learning (RL) in episodic Markov decision processes, where experience from related source tasks is available during learning on a target task. A fundamental difficulty is that task similarity is…
You measure the value of a quantity x for a number of systems (cells, molecules, people, chunks of metal, DNA vectors, etc.). You repeat the whole set of measures in different occasions or assays, which you try to design as equal to one…
There has been growing progress on theoretical analyses for provably efficient learning in MDPs with linear function approximation, but much of the existing work has made strong assumptions to enable exploration by conventional exploration…
The paper is about developing a solver for maximizing a real-valued function of binary variables. The solver relies on an algorithm that estimates the optimal objective-function value of instances from the underlying distribution of…
This article contains a review of Nelson's analysis of Bell's theorem. It shows that Bell's inequalities can be violated with a theory of local random variables if one accepts that the outcomes of these variables are not predetermined prior…
We consider retarded settings in the context of a Bell-type experiment. The retarded setting is defined as the value the setting would have taken were it not for some external intervention (for example, by a human). We derive retarded Bell…
Large language models (LLMs) are increasingly being used for tasks where outputs shape human decisions, so it is critical to verify that their responses consistently reflect desired human values. Humans, as individuals or groups, don't…
Benchmarks underpin how progress in large language models (LLMs) is measured and trusted. Yet our analyses reveal that apparent convergence in benchmark accuracy can conceal deep epistemic divergence. Using two major reasoning benchmarks -…
We study an optimal control problem in Bolza form and we consider the value function associated to this problem. We prove two verification theorems which ensure that, if a function $W$ satisfies some suitable weak continuity assumptions and…