中文
相关论文

相关论文: Doubly-Asynchronous Value Iteration: Making Value …

200 篇论文

This study presents the vectorization of metaheuristic algorithms as the first stage of vectorized optimization implementation. Vectorization is a technique for converting an algorithm, which operates on a single value at a time to one that…

神经与进化计算 · 计算机科学 2023-08-22 Mahmood Yashar , Tarik A. Rashid

We develop a parameterized Primal-Dual $\pi$ Learning method based on deep neural networks for Markov decision process with large state space and off-policy reinforcement learning. In contrast to the popular Q-learning and actor-critic…

机器学习 · 计算机科学 2017-12-08 Woon Sang Cho , Mengdi Wang

This paper considers online optimization for a system that performs a sequence of back-to-back tasks. Each task can be processed in one of multiple processing modes that affect the duration of the task, the reward earned, and an additional…

最优化与控制 · 数学 2024-01-17 Michael J. Neely

We propose universal randomized function approximation-based empirical value iteration (EVI) algorithms for Markov decision processes. The `empirical' nature comes from each iteration being done empirically from samples available from…

最优化与控制 · 数学 2019-04-25 William B. Haskell , Rahul Jain , Hiteshi Sharma , Pengqian Yu

We present DeepMVI, a deep learning method for missing value imputation in multidimensional time-series datasets. Missing values are commonplace in decision support platforms that aggregate data over long time stretches from disparate…

机器学习 · 计算机科学 2023-06-22 Parikshit Bansal , Prathamesh Deshpande , Sunita Sarawagi

In this paper, we develop a new accelerated stochastic gradient method for efficiently solving the convex regularized empirical risk minimization problem in mini-batch settings. The use of mini-batches is becoming a golden standard in the…

最优化与控制 · 数学 2017-09-20 Tomoya Murata , Taiji Suzuki

An asynchronous, variational method for simulating elastica in complex contact and impact scenarios is developed. Asynchronous Variational Integrators (AVIs) are extended to handle contact forces by associating different time steps to…

数值分析 · 数学 2015-05-19 Etienne Vouga , David Harmon , Rasmus Tamstorf , Eitan Grinspun

We present a parallelized primal-dual algorithm for solving constrained convex optimization problems. The algorithm is "block-based," in that vectors of primal and dual variables are partitioned into blocks, each of which is updated only by…

最优化与控制 · 数学 2020-09-01 Katherine Hendrickson , Matthew Hale

This note provides a simple example demonstrating that, if exact computations are allowed, the number of iterations required for the value iteration algorithm to find an optimal policy for discounted dynamic programming problems may grow…

人工智能 · 计算机科学 2013-12-25 Eugene A. Feinberg , Jefferson Huang

We build on a recently introduced geometric interpretation of Markov Decision Processes (MDPs) to analyze classical MDP-solving algorithms: Value Iteration (VI) and Policy Iteration (PI). First, we develop a geometry-based analytical…

机器学习 · 计算机科学 2025-03-07 Arsenii Mustafin , Aleksei Pakharev , Alex Olshevsky , Ioannis Ch. Paschalidis

The Value Iteration Network (VIN) is an end-to-end differentiable neural network architecture for planning. It exhibits strong generalization to unseen domains by incorporating a differentiable planning module that operates on a latent…

机器学习 · 计算机科学 2025-07-08 Yuhui Wang , Qingyuan Wu , Dylan R. Ashley , Francesco Faccio , Weida Li , Chao Huang , Jürgen Schmidhuber

The proliferation of computing devices has brought about an opportunity to deploy machine learning models on new problem domains using previously inaccessible data. Traditional algorithms for training such models often require data to be…

Recent non-asymptotic analyses have substantially advanced the theory of distributional policy evaluation, but they largely concern synchronous full-state updates under a generative model, model-based estimators, accelerated variants, or…

机器学习 · 计算机科学 2026-05-11 Ege C. Kaya , Abolfazl Hashemi

Deep learning has revolutionized the last decade, being at the forefront of extraordinary advances in a wide range of tasks including computer vision, natural language processing, and reinforcement learning, to name but a few. However, it…

机器学习 · 计算机科学 2024-01-24 Sebastian W. Ober

Stochastic variational inference (SVI) employs stochastic optimization to scale up Bayesian computation to massive data. Since SVI is at its core a stochastic gradient-based algorithm, horizontal parallelism can be harnessed to allow larger…

机器学习 · 统计学 2018-01-16 Saad Mohamad , Abdelhamid Bouchachia , Moamar Sayed-Mouchaweh

We introduce a continuous policy-value iteration algorithm where the approximations of the value function of a stochastic control problem and the optimal control are simultaneously updated through Langevin-type dynamics. This framework…

最优化与控制 · 数学 2025-06-11 Qi Feng , Gu Wang

Autonomous robot manipulation is a complex and continuously evolving robotics field. This paper focuses on data augmentation methods in imitation learning. Imitation learning consists of three stages: data collection from experts, learning…

机器人学 · 计算机科学 2024-10-08 Masato Kobayashi , Thanpimon Buamanee , Yuki Uranishi

Variational empirical Bayes (VEB) methods provide a practically attractive approach to fitting large, sparse, multiple regression models. These methods usually use coordinate ascent to optimize the variational objective function, an…

统计方法学 · 统计学 2024-11-25 Saikat Banerjee , Peter Carbonetto , Matthew Stephens

We present a method of exploiting symmetries of discrete-time optimal control problems to reduce the dimensionality of dynamic programming iterations. The results are derived for systems with continuous state variables, and can be applied…

系统与控制 · 计算机科学 2018-08-30 John Maidens , Axel Barrau , Silvere Bonnabel , Murat Arcak

This paper considers optimization over multiple renewal systems coupled by time average constraints. These systems act asynchronously over variable length frames. For each system, at the beginning of each renewal frame, it chooses an action…

最优化与控制 · 数学 2018-05-23 Xiaohan Wei , Michael J. Neely