Related papers: The Target Discounted-Sum Problem
In recent years there has been intense interest in the vanishing discount problem for Hamilton-Jacobi equations. In the case of the scalar equation, B. Ziliotto has recently given an example of the Hamilton-Jacobi equation having non-convex…
This paper studies semiparametric contextual bandits, a generalization of the linear stochastic bandit problem where the reward for an action is modeled as a linear function of known action features confounded by an non-linear…
The Promise Constraint Satisfaction Problem (PCSP) is a generalization of the Constraint Satisfaction Problem (CSP) that includes approximation variants of satisfiability and graph coloring problems. Barto [LICS '19] has shown that a…
While discounted payoff games and classic games that reduce to them, like parity and mean-payoff games, are symmetric, their solutions are not. We have taken a fresh view on the properties that optimal solutions need to have, and devised a…
While discounted payoff games and classic games that reduce to them, like parity and mean-payoff games, are symmetric, their solutions are not. We have taken a fresh view on the constraints that optimal solutions need to satisfy, and…
In this article we propose a Weighted Stochastic Mesh (WSM) Algorithm for approximating the value of a discrete and continuous time optimal stopping problem. We prove that in the discrete case the WSM algorithm leads to semi-tractability of…
Let $(\xi_1, \eta_1)$, $(\xi_2, \eta_2),\ldots$ be independent identically distributed $\mathbb{R}^2$-valued random vectors. Assuming that $\xi_1$ has zero mean and finite variance and imposing three distinct groups of assumptions on the…
The stochastic composition optimization proposed recently by Wang et al. [2014] minimizes the objective with the compositional expectation form: $\min_x~(\mathbb{E}_iF_i \circ \mathbb{E}_j G_j)(x).$ It summarizes many important applications…
In this paper we consider a fragment of the first-order theory of the real numbers that includes systems of equations of continuous functions in bounded domains, and for which all functions are computable in the sense that it is possible to…
Let $f\in \mathbb{Q}(x)$ be a non-constant rational function. We consider "Waring's Problem for $f(x)$," i.e., whether every element of $\bbq$ can be written as a bounded sum of elements of $\{f(a)\mid a\in \mathbb{Q}\}$. For rational…
We consider optimization problems in which the objective requires an inner loop with many steps or is the limit of a sequence of increasingly costly approximations. Meta-learning, training recurrent neural networks, and optimization of the…
We consider a budget-constrained bandit problem where each arm pull incurs a random cost, and yields a random reward in return. The objective is to maximize the total expected reward under a budget constraint on the total cost. The model is…
We study the links between the values of stochastic games with varying stage duration $h$, the corresponding Shapley operators $\bf{T}$ and ${\bf{T}}\_h$and the solution of $\dot f\_t = ({\bf{T}} - Id )f\_t$. Considering general non…
Given two weighted automata, we consider the problem of whether one is big-O of the other, i.e., if the weight of every finite word in the first is not greater than some constant multiple of the weight in the second. We show that the…
In this paper we establish a new summation method by expanding $\prod_{k}(1-\frac{z}{a_{k}})^{-1}$ with two approaches: the Taylor expansion and the infinite partial fraction decomposition. Here we focus on the case when $a_{k}$ is…
We consider the two categories of termination problems of quantum programs with nondeterminism: 1) Is an input of a program terminating with probability one under all schedulers? If not, how can a scheduler be synthesized to evidence the…
We consider an adversarial variant of the classic $K$-armed linear contextual bandit problem where the sequence of loss functions associated with each arm are allowed to change without restriction over time. Under the assumption that the…
We take a unifying approach to single selection optimal stopping problems with random arrival order and independent sampling of items. In the problem we consider, a decision maker (DM) initially gets to sample each of $N$ items…
We design alignment-free techniques for comparing a sequence or word, called a target, against a set of words, called a reference. A target-specific factor of a target $T$ against a reference $R$ is a factor $w$ of a word in $T$ which is…
We describe a novel optimization method for finite sums (such as empirical risk minimization problems) building on the recently introduced SAGA method. Our method achieves an accelerated convergence rate on strongly convex smooth problems.…