Related papers: A New Approach to Drifting Games, Based on Asympto…
We use optimism to introduce generic asymptotically optimal reinforcement learning agents. They achieve, with an arbitrary finite or compact class of environments, asymptotically optimal behavior. Furthermore, in the finite deterministic…
Drifting is a complicated task for autonomous vehicle control. Most traditional methods in this area are based on motion equations derived by the understanding of vehicle dynamics, which is difficult to be modeled precisely. We propose a…
In this paper we provide a thorough, rigorous theoretical framework to assess optimality guarantees of sampling-based algorithms for drift control systems: systems that, loosely speaking, can not stop instantaneously due to momentum. We…
We provide a decision theoretic analysis of bandit experiments under local asymptotics. Working within the framework of diffusion processes, we define suitable notions of asymptotic Bayes and minimax risk for these experiments. For normally…
The main contributions of this paper are three fold. First, our primary concern is to investigate a class of stochastic recursive delayed control problems which arise naturally with sound backgrounds but have not been well-studied yet. For…
In this paper, we study the asymptotic behavior as $x_1\to+\infty$ of solutions of semilinear elliptic equations in quarter- or half-spaces, for which the value at $x_1=0$ is given. We prove the uniqueness and characterize the…
In the contextual linear bandit setting, algorithms built on the optimism principle fail to exploit the structure of the problem and have been shown to be asymptotically suboptimal. In this paper, we follow recent approaches of deriving…
We study a class of reflected backward stochastic differential equations with nonpositive jumps and upper barrier. Existence and uniqueness of a minimal solution is proved by a double penalization approach under regularity assumptions on…
The paper is concerned with two-person games with saddle point. We investigate the limits of value functions for long-time-average payoff, discounted average payoff, and the payoff that follows a probability density. Most of our assumptions…
Static potential games are non-cooperative games which admit a fictitious function, also referred to as a potential function, such that the minimizers of this function constitute a subset (or a refinement) of the Nash equilibrium strategies…
In this work, we establish near-linear and strong convergence for a natural first-order iterative algorithm that simulates Von Neumann's Alternating Projections method in zero-sum games. First, we provide a precise analysis of Optimistic…
The dominant line of work in domain adaptation has focused on learning invariant representations using domain-adversarial training. In this paper, we interpret this approach from a game theoretical perspective. Defining optimal solutions in…
The focus of this paper is to propose a driver model that incorporates human reasoning levels as actions during interactions with other drivers. Different from earlier work using game theoretical human reasoning levels, we propose a dynamic…
We provide several applications of Optimistic Mirror Descent, an online learning algorithm based on the idea of predictable sequences. First, we recover the Mirror Prox algorithm for offline optimization, prove an extension to Holder-smooth…
We consider a sequence of repeated prediction games and formally pass to the limit. The supersolutions of the resulting non-linear parabolic partial differential equation are closely related to the potential functions in the sense of…
Traditional solvable game theory and mean-field-type game theory (risk-aware games) predominantly focus on quadratic costs due to their analytical tractability. Nevertheless, they often fail to capture critical non-linearities inherent in…
This paper proposes a payoff perturbation technique for the Mirror Descent (MD) algorithm in games where the gradient of the payoff functions is monotone in the strategy profile space, potentially containing additive noise. The optimistic…
In a noncooperative dynamic game, multiple agents operating in a changing environment aim to optimize their utilities over an infinite time horizon. Time-varying environments allow to model more realistic scenarios (e.g., mobile devices…
Lagrangian motions of fluid particles in a general velocity field oscillating in time are studied with the use of the two-timing method. Our aims are: (i) to calculate systematically the most general and practically usable asymptotic…
We provide a deterministic-control-based interpretation for a broad class of fully nonlinear parabolic and elliptic PDEs with continuous Neumann boundary conditions in a smooth domain. We construct families of two-person games depending on…