相关论文: (ANTI)PETER Principle - Discrete (INVERSE) Logisti…
The challenge of mastering computational tasks of enormous size tends to frequently override questioning the quality of the numerical outcome in terms of accuracy. By this we do not mean the accuracy within the discrete setting, which…
Modeling the purposeful behavior of imperfect agents from a small number of observations is a challenging task. When restricted to the single-agent decision-theoretic setting, inverse optimal control techniques assume that observed behavior…
The policy relevant treatment effect (PRTE) measures the average effect of switching from a status-quo policy to a counterfactual policy. Estimation of the PRTE involves estimation of multiple preliminary parameters, including propensity…
We study the limit behaviour of upper and lower bounds on expected time averages in imprecise Markov chains; a generalised type of Markov chain where the local dynamics, traditionally characterised by transition probabilities, are now…
Learning conditional distributions $\pi^*(\cdot|x)$ is a central problem in machine learning, which is typically approached via supervised methods with paired data $(x,y) \sim \pi^*$. However, acquiring paired data samples is often…
In this paper, we consider the problem of learning safe policies for probabilistic-constrained reinforcement learning (RL). Specifically, a safe policy or controller is one that, with high probability, maintains the trajectory of the agent…
Complex planning and scheduling problems have long been solved using various optimization or heuristic approaches. In recent years, imitation learning that aims to learn from expert demonstrations has been proposed as a viable alternative…
Inverse optimal control can be used to characterize behavior in sequential decision-making tasks. Most existing work, however, is limited to fully observable or linear systems, or requires the action signals to be known. Here, we introduce…
We consider the variational discretization of a linear-quadratic optimal control problem with pointwise control and state constraints. In order to allow for a Fr\'echet smooth norm, the problem is reformulated by means of a reflexive…
Inverse reinforcement learning is the problem of inferring a reward function from an optimal policy or demonstrations by an expert. In this work, it is assumed that the reward is expressed as a reward machine whose transitions depend on…
In this paper, we consider a hierarchical control problem with model uncertainty. Specifically, we consider the following objectives that we would like to accomplish. The first one being of a controllability-type that consists of…
In many classification settings, the class of primary interest is underrepresented, leading to imbalanced data problems that arise in applications such as rare disease detection and fraud identification. In these contexts, identifying a…
In many choice settings the decision maker (DM) adopts a criterion which is a mediation between her preference, and its opposite. According to such compromise, the first i alternatives on top of the DM's taste are moved, in reverse order,…
We prove a number of \textit{a priori} estimates for weak solutions of elliptic equations or systems with vertically independent coefficients in the upper-half space. These estimates are designed towards applications to boundary value…
This paper presents an approach for developing the explanation capabilities of rule-based expert systems managing imprecise and uncertain knowledge. The treatment of uncertainty takes place in the framework of possibility theory where the…
Reinforcement learning for large language models faces a fundamental trade-off between sample efficiency and asymptotic performance: strictly on-policy methods discard trajectories after a single update, while off-policy reuse introduces…
This work introduces a model in which agents of a network act upon one another according to three different kinds of moral decisions. These decisions are based on an increasing level of sophistication in the empathy capacity of the agent, a…
The double descent (DD) paradox, where over-parameterized models see generalization improve past the interpolation point, remains largely unexplored in the non-stationary domain of Deep Reinforcement Learning (DRL). We present preliminary…
The St. Petersburg paradox presents a longstanding challenge in decision theory: its classical expected value diverges, yet no correspondingly large finite stake is typically regarded as rational. Traditional responses introduce auxiliary…
In this note, we propose a discrete model to study one-dimensional transport equations with non-local drift and supercritical dissipation. The inspiration for our model is the equation $$ \theta_t + (H\theta) \theta_x +(-\Delta)^\alpha…