English
Related papers

Related papers: Relative Divergence and Maximum Relative Divergenc…

200 papers

The concept of Relative Divergence of one Grading Function from another is extended from totally ordered chains to power sets of finite event spaces. Shannon Entropy concept is extended to normalized grading functions on such power sets.…

Probability · Mathematics 2022-07-15 Alexander Dukhovny

The concept of Shannon Entropy for probability distributions and associated Maximum Entropy Principle are extended here to the concepts of Relative Divergence of one Grading Function from another and Maximum Relative Divergence Principle…

Optimization and Control · Mathematics 2023-03-28 Alexander Dukhovny

The standard conditional probability definition formula is derived as a consequence of the Insufficient Reason Principle expressed as the Maximum Relative Divergence Principle for grading (order-comonotonic) functions on a totally ordered…

Probability · Mathematics 2025-07-14 Alexander Dukhovny

Robust Markov decision processes (RMDPs) extend standard Markov decision processes (MDPs) to account for uncertainty in the transition probabilities. RMDPs have an uncertainty set that defines a set of possible transition functions, each of…

Logic in Computer Science · Computer Science 2026-04-30 Marnix Suilen , Guillermo A. Pérez

Fueled by advances in both robust optimization theory and reinforcement learning (RL), robust Markov Decision Processes (RMDPs) have garnered increasing attention due to their powerful capability for sequential decision-making under…

Optimization and Control · Mathematics 2025-07-08 Wenfan Ou , Sheng Bi

We address the problem of computing reliable policies in reinforcement learning problems with limited data. In particular, we compute policies that achieve good returns with high confidence when deployed. This objective, known as the…

Machine Learning · Computer Science 2021-03-01 Bahram Behzadian , Reazul Hasan Russel , Marek Petrik , Chin Pang Ho

By comparing the original equations with the corresponding stationary ones, the moderate deviation principle (MDP) is established for unbounded additive functionals of several different models of distribution dependent SDEs, with…

Probability · Mathematics 2021-01-26 Panpan Ren , Shen Wang

Robust Markov decision processes (RMDPs) provide a promising framework for computing reliable policies in the face of model errors. Many successful reinforcement learning algorithms build on variations of policy-gradient methods, but…

Machine Learning · Computer Science 2024-05-15 Qiuhao Wang , Chin Pang Ho , Marek Petrik

The standard definition formula for probabilities of independent events is derived as a consequence of the Insufficient Reason Principle expressed as the Maximum Relative Divergence Principle for grading (order-comonotonic) functions on a…

Probability · Mathematics 2025-08-05 Alexander Dukhovny

We consider Markov decision processes (MDPs) with unknown disturbance distribution and address this problem using the robust Markov decision process (RMDP) approach. We construct the empirical distribution of the unknown disturbance…

Optimization and Control · Mathematics 2026-03-11 Sivaramakrishnan Ramani

One key challenge for multi-task Reinforcement learning (RL) in practice is the absence of task indicators. Robust RL has been applied to deal with task ambiguity, but may result in over-conservative policies. To balance the worst-case…

Machine Learning · Computer Science 2022-10-25 Mengdi Xu , Peide Huang , Yaru Niu , Visak Kumar , Jielin Qiu , Chao Fang , Kuan-Hui Lee , Xuewei Qi , Henry Lam , Bo Li , Ding Zhao

In this paper we analyze the joint rate distortion function (RDF), for a tuple of correlated sources taking values in abstract alphabet spaces (i.e., continuous) subject to two individual distortion criteria. First, we derive structural…

Information Theory · Computer Science 2021-05-11 Evagoras Stylianou , Charalambos D. Charalambous , Themistoklis Charalambous

Policy gradient methods are among the most effective methods in challenging reinforcement learning problems with large state and/or action spaces. However, little is known about even their most basic theoretical convergence properties,…

Machine Learning · Computer Science 2020-10-16 Alekh Agarwal , Sham M. Kakade , Jason D. Lee , Gaurav Mahajan

The Robust Markov Decision Process (RMDP) framework focuses on designing control policies that are robust against the parameter uncertainties due to the mismatches between the simulator model and real-world settings. An RMDP problem is…

Machine Learning · Computer Science 2022-05-17 Kishan Panaganti , Dileep Kalathil

Recent advances in Rate-Distortion-Perception (RDP) theory highlight the importance of balancing compression level, reconstruction quality, and perceptual fidelity. While previous work has explored numerical approaches to approximate the…

Information Theory · Computer Science 2025-08-20 Chunhui Chen , Linyi Chen , Xueyan Niu , Hao Wu

The joint replenishment problem (JRP) is a classical inventory management problem. We consider a natural generalization with outliers, where we are allowed to reject (that is, not service) a subset of demand points. In this paper, we are…

Data Structures and Algorithms · Computer Science 2023-08-10 Varun Suriyanarayana , Varun Sivashankar , Siddharth Gollapudi , David Shmoys

Random permutation set (RPS) is a new formalism for reasoning with uncertainty involving order information. Measuring the conflict between two pieces of evidence represented by permutation mass functions remains an open issue in…

Artificial Intelligence · Computer Science 2026-03-20 Ruolan Cheng , Yong Deng

Robust MDPs (RMDPs) can be used to compute policies with provable worst-case guarantees in reinforcement learning. The quality and robustness of an RMDP solution are determined by the ambiguity set---the set of plausible transition…

Machine Learning · Computer Science 2019-02-21 Marek Petrik , Reazul Hasan Russell

This paper studies the rate-distortion-perception (RDP) tradeoff for a memoryless source model in the asymptotic limit of large block-lengths. The perception measure is based on a divergence between the distributions of the source and…

Information Theory · Computer Science 2025-04-29 Sadaf Salehkalaibar , Jun Chen , Ashish Khisti , Wei Yu

The Probability Ranking Principle (PRP) has been considered as the foundational standard in the design of information retrieval (IR) systems. The principle requires an IR module's returned list of results to be ranked with respect to the…

Information Retrieval · Computer Science 2024-05-09 Kai Zheng , Haijun Zhao , Rui Huang , Beichuan Zhang , Na Mou , Yanan Niu , Yang Song , Hongning Wang , Kun Gai
‹ Prev 1 2 3 10 Next ›