English
Related papers

Related papers: Relative Divergence and Maximum Relative Divergenc…

200 papers

Shuffling gradient methods are widely used in modern machine learning tasks and include three popular implementations: Random Reshuffle (RR), Shuffle Once (SO), and Incremental Gradient (IG). Compared to the empirical success, the…

Machine Learning · Computer Science 2024-06-07 Zijian Liu , Zhengyuan Zhou

Reward models (RMs) are central to aligning large language models, yet their practical effectiveness hinges on generalization to unseen prompts and shifting distributions. Most existing RM evaluations rely on static, pre-annotated…

Computation and Language · Computer Science 2026-01-27 Shunyang Luo , Peibei Cao , Zhihui Zhu , Kehua Feng , Zhihua Wang , Keyan Ding

The pursuit of robustness has recently been a popular topic in reinforcement learning (RL) research, yet the existing methods generally suffer from efficiency issues that obstruct their real-world implementation. In this paper, we introduce…

Machine Learning · Computer Science 2024-04-15 Yang Hu , Haitong Ma , Bo Dai , Na Li

In real life, we frequently come across data sets that involve some independent explanatory variable(s) generating a set of ordinal responses. These ordinal responses may correspond to an underlying continuous latent variable, which is…

Methodology · Statistics 2024-01-08 Arijit Pyne , Subhrajyoty Roy , Abhik Ghosh , Ayanendranath Basu

Markov decision processes (MDPs) are a popular model for performance analysis and optimization of stochastic systems. The parameters of stochastic behavior of MDPs are estimates from empirical observations of a system; their values are not…

Artificial Intelligence · Computer Science 2017-10-26 Dimitri Scheftelowitsch , Peter Buchholz , Vahid Hashemi , Holger Hermanns

The Moderate Deviations Principle (MDP) is well-understood for sums of independent random variables, worse understood for stationary random sequences, and scantily understood for random fields. Here it is established for some planary random…

Probability · Mathematics 2019-09-25 Boris Tsirelson

The distributionally robust Markov Decision Process (MDP) approach asks for a distributionally robust policy that achieves the maximal expected total reward under the most adversarial distribution of uncertain parameters. In this paper, we…

Systems and Control · Computer Science 2018-10-10 Zhi Chen , Pengqian Yu , William B. Haskell

The $\chi^2$-principle generalizes the Morozov discrepancy principle (MDP) to the augmented residual of the Tikhonov regularized least squares problem. Weighting of the data fidelity by a known Gaussian noise distribution on the measured…

Numerical Analysis · Mathematics 2022-08-16 Saeed Vatankhah , Rosemary A Renaut , Vahid E Ardestani

Marginal maximum likelihood estimation (MMLE) in item response theory (IRT) is highly sensitive to aberrant responses, such as careless answering and random guessing, which can reduce estimation accuracy. To address this issue, this study…

Methodology · Statistics 2025-02-18 Yuki Itaya , Kenichi Hayashi

In online Inverse Reinforcement Learning (IRL), the learner can collect samples about the dynamics of the environment to improve its estimate of the reward function. Since IRL suffers from identifiability issues, many theoretical works on…

Machine Learning · Computer Science 2024-10-10 Filippo Lazzati , Mirco Mutti , Alberto Maria Metelli

We consider the reinforcement learning problem for the constrained Markov decision process (CMDP), which plays a central role in satisfying safety or resource constraints in sequential learning and decision-making. In this problem, we are…

Machine Learning · Computer Science 2025-11-19 Jiashuo Jiang , Yinyu Ye

The problem of estimating the information rate distortion perception function (RDPF), which is a relevant information-theoretic quantity in goal-oriented lossy compression and semantic information reconstruction, is investigated here.…

Information Theory · Computer Science 2025-09-25 Martha V. Sourla , Giuseppe Serra , Photios A. Stavrou , Marios Kountouris

Robustness is important for sequential decision making in a stochastic dynamic environment with uncertain probabilistic parameters. We address the problem of using robust MDPs (RMDPs) to compute policies with provable worst-case guarantees…

Machine Learning · Computer Science 2018-11-16 Reazul Hasan Russel , Marek Petrik

It is well-understood that different algorithms, training processes, and corpora produce different word embeddings. However, less is known about the relation between different embedding spaces, i.e. how far different sets of embeddings…

Computation and Language · Computer Science 2020-05-19 Xuhui Zhou , Zaixiang Zheng , Shujian Huang

We consider undiscounted reinforcement learning in Markov decision processes (MDPs) where both the reward functions and the state-transition probabilities may vary (gradually or abruptly) over time. For this problem setting, we propose an…

Machine Learning · Computer Science 2019-09-11 Pratik Gajane , Ronald Ortner , Peter Auer

We study the problem of inverse reinforcement learning (IRL), where the learning agent recovers a reward function using expert demonstrations. Most of the existing IRL techniques make the often unrealistic assumption that the agent has…

Machine Learning · Computer Science 2021-12-20 Franck Djeumou , Murat Cubuktepe , Craig Lennon , Ufuk Topcu

Robust reinforcement learning is essential for deploying reinforcement learning algorithms in real-world scenarios where environmental uncertainty predominates. Traditional robust reinforcement learning often depends on rectangularity…

Machine Learning · Computer Science 2024-06-13 Adil Zouitine , David Bertoin , Pierre Clavier , Matthieu Geist , Emmanuel Rachelson

We study a generic class of \emph{random optimization problems} (rops) and their typical behavior. The foundational aspects of the random duality theory (RDT), associated with rops, were discussed in \cite{StojnicRegRndDlt10}, where it was…

Probability · Mathematics 2023-12-04 Mihailo Stojnic

Inverse reinforcement learning is the problem of inferring a reward function from an optimal policy or demonstrations by an expert. In this work, it is assumed that the reward is expressed as a reward machine whose transitions depend on…

Machine Learning · Computer Science 2025-10-23 Mohamad Louai Shehab , Antoine Aspeel , Necmiye Ozay

Generalization of the Lorden's inequality is an excellent tool for obtaining strong upper bounds for the convergence rate for various complicated stochastic models. This paper demonstrates a method for obtaining such bounds for some…

Probability · Mathematics 2020-10-13 Galina Zverkina
‹ Prev 1 4 5 6 7 8 10 Next ›