中文
相关论文

相关论文: Approximate Dynamic Programming based on High Dime…

200 篇论文

Approximate linear programming (ALP) is an efficient approach to solving large factored Markov decision processes (MDPs). The main idea of the method is to approximate the optimal value function by a set of basis functions and optimize…

人工智能 · 计算机科学 2012-06-18 Branislav Kveton , Milos Hauskrecht

This paper addresses a fundamental issue central to approximation methods for solving large Markov decision processes (MDPs): how to automatically learn the underlying representation for value function approximation? A novel theoretically…

人工智能 · 计算机科学 2012-07-09 Sridhar Mahadevan

In this note, we develop fast and deterministic dimensionality reduction techniques for a family of subspace approximation problems. Let $P\subset \mathbbm{R}^N$ be a given set of $M$ points. The techniques developed herein find an $O(n…

计算几何 · 计算机科学 2013-12-06 Mark Iwen , Felix Krahmer

This paper provides an approximate online adaptive solution to the infinite-horizon optimal tracking problem for control-affine continuous-time nonlinear systems with unknown drift dynamics. Model-based reinforcement learning is used to…

系统与控制 · 计算机科学 2017-07-25 Rushikesh Kamalapurkar , Lindsey Andrews , Patrick Walters , Warren E. Dixon

In this paper we consider the numerical approximation of infinite horizon problems via the dynamic programming approach. The value function of the problem solves a Hamilton-Jacobi-Bellman (HJB) equation that is approximated by a fully…

数值分析 · 数学 2024-11-06 Javier de Frutos , Bosco Garcia-Archilla , Julia Novo

Deep Reinforcement Learning has shown its ability in solving complicated problems directly from high-dimensional observations. However, in end-to-end settings, Reinforcement Learning algorithms are not sample-efficient and requires long…

机器学习 · 计算机科学 2021-07-06 Nicolò Botteghi , Mannes Poel , Beril Sirmacek , Christoph Brune

For safely applying reinforcement learning algorithms on high-dimensional nonlinear dynamical systems, a simplified system model is used to formulate a safe reinforcement learning framework. Based on the simplified system model, a…

机器人学 · 计算机科学 2021-09-09 Zhehua Zhou , Ozgur S. Oguz , Marion Leibold , Martin Buss

In real-world applications, it is important for machine learning algorithms to be robust against data outliers or corruptions. In this paper, we focus on improving the robustness of a large class of learning algorithms that are formulated…

机器学习 · 计算机科学 2021-06-04 Quanming Yao , Hangsi Yang , En-Liang Hu , James Kwok

We consider the problem of minimizing a proper, lower semicontinuous, geodesically convex function on a Hadamard manifold. Building on ball-proximal (broximal) ideas in the Euclidean setting, viewed as an abstract proximal-type algorithm,…

最优化与控制 · 数学 2026-05-06 F. Babu , O. P. Ferreira , L. F. Prudente , Jen-Chih Yao , Xiaopeng Zhao

High-dimensional models often have a large memory footprint and must be quantized after training before being deployed on resource-constrained edge devices for inference tasks. In this work, we develop an information-theoretic framework for…

信息论 · 计算机科学 2022-09-01 Rajarshi Saha , Mert Pilanci , Andrea J. Goldsmith

Hidden semi-Markov models (HSMMs) are latent variable models which allow latent state persistence and can be viewed as a generalization of the popular hidden Markov models (HMMs). In this paper, we introduce a novel spectral algorithm to…

机器学习 · 统计学 2016-03-01 Igor Melnyk , Arindam Banerjee

This paper addresses the problem of finding a B-term wavelet representation of a given discrete function $f \in \real^n$ whose distance from f is minimized. The problem is well understood when we seek to minimize the Euclidean distance…

数据结构与算法 · 计算机科学 2007-07-23 Sudipto Guha , Boulos Harb

We consider dynamic programming problems with finite, discrete-time horizons and prohibitively high-dimensional, discrete state-spaces for direct computation of the value function from the Bellman equation. For the case that the value…

最优化与控制 · 数学 2020-05-25 Denis Lebedev , Paul Goulart , Kostas Margellos

Random projection (RP) is a classical technique for reducing storage and computational costs. We analyze RP-based approximations of convex programs, in which the original optimization problem is approximated by the solution of a…

信息论 · 计算机科学 2014-04-30 Mert Pilanci , Martin J. Wainwright

This paper introduces a fast algorithm for randomized computation of a low-rank Dynamic Mode Decomposition (DMD) of a matrix. Here we consider this matrix to represent the development of a spatial grid through time e.g. data from a static…

计算机视觉与模式识别 · 计算机科学 2016-04-12 N. Benjamin Erichson , Carl Donovan

Low-rank approximation of a matrix by means of structured random sampling has been consistently efficient in its extensive empirical studies around the globe, but adequate formal support for this empirical phenomenon has been missing so…

数值分析 · 数学 2016-07-21 Victor Pan , John Svadlenka , Liang Zhao

The MM principle is a device for creating optimization algorithms satisfying the ascent or descent property. The current survey emphasizes the role of the MM principle in nonlinear programming. For smooth functions, one can construct an…

最优化与控制 · 数学 2015-07-29 Kenneth Lange , Kevin L. Keys

Value function approximation is important in modern reinforcement learning (RL) problems especially when the state space is (infinitely) large. Despite the importance and wide applicability of value function approximation, its theoretical…

机器学习 · 计算机科学 2023-02-24 Hanlin Zhu , Ruosong Wang , Jason D. Lee

Dynamic mode decomposition (DMD) is a widely used data-driven algorithm for predicting the future states of dynamical systems. However, its standard formulation often struggles with poor long-term predictive accuracy. To address this…

数值分析 · 数学 2025-10-23 Qiuqi Li , Chang Liu , Yifei Yang

In the paper, we consider the problem of robust approximation of transfer Koopman and Perron-Frobenius (P-F) operators from noisy time series data. In most applications, the time-series data obtained from simulation or experiment is…

最优化与控制 · 数学 2020-01-08 Subhrajit Sinha , Huang Bowen , Umesh Vaidya
‹ 上一页 1 8 9 10 下一页 ›