Related papers: MRL order, log-concavity and an application to pea…
We study a class of stationary Markov processes with marginal distributions identifiable by moments such that every conditional moment of degree say $m$ is a polynomial of degree at most $m\;\text{.}\;$ We show that then under some…
We provide a new characterization of second-order stochastic dominance, also known as increasing concave order. The result has an intuitive interpretation that adding a risk with negative expected value in adverse scenarios makes the…
For a countable-state Markov decision process we introduce an embedding which produces a finite-state Markov decision process. The finite-state embedded process has the same optimal cost, and moreover, it has the same dynamics as the…
Model-based reinforcement learning (MBRL) is a crucial approach to enhance the generalization capabilities and improve the sample efficiency of RL algorithms. However, current MBRL methods focus primarily on building world models for single…
In this paper the class of mixed renewal processes (MRPs for short) with mixing parameter a random vector from \cite{lm6z3} (enlarging Huang's \cite{hu} original class) is replaced by the strictly more comprising class of all extended MRPs…
Examples of stochastic processes whose state space representations involve functions of an integral type structure $$I_{t}^{(a,b)}:=\int_{0}^{t}b(Y_{s})e^{-\int_{s}^{t}a(Y_{r})dr}ds, \quad t\ge 0$$ are studied under an ergodic…
Suppose that $X_1, \ldots , X_n$ are continuous semimartingales that are reversible and have nondegenerate crossings. Then the corresponding rank processes can be represented by generalized Stratonovich integrals, and this representation…
This paper explores some sufficient conditions for the enhanced solvability of strong vector equilibrium problems, which can be established via a variational approach. Enhanced solvability here means existence of solutions, which are strong…
This paper develops an axiomatic framework for ranking metrics, a general class of functionals for evaluating and ordering financial or insurance positions. Unlike traditional risk-adjusted performance measures-such as the Sharpe ratio,…
Geometric properties can be leveraged to stabilize and speed reinforcement learning. Existing examples include encoding symmetry structure, geometry-aware data augmentation, and enforcing structural restrictions. In this paper, we take a…
We introduce Prompt Curriculum Learning (PCL), a lightweight reinforcement learning (RL) algorithm that selects intermediate-difficulty prompts using a learned value model to post-train language models. Since post-training LLMs via RL…
In this paper we propose an open anomalous semi-Markovian random neural networks model with negative and positive signals with arbitrary random waiting times. We investigate the signal flow process in the anomalous random neural networks…
We introduce a self-reinforced point processes on the unit interval that appears to exhibit self-organized criticality, somewhat reminiscent of the well-known Bak-Sneppen model. The process takes values in the finite subsets of the unit…
In this paper stochastic partitioned Runge-Kutta (SPRK) methods are considered. A general order theory for SPRK methods based on stochastic B-series and multicolored, multishaped rooted trees is developed. The theory is applied to prove the…
Deep learning has been shown to achieve impressive results in several tasks where a large amount of training data is available. However, deep learning solely focuses on the accuracy of the predictions, neglecting the reasoning process…
Evolutionary algorithms (EA), a class of stochastic search methods based on the principles of natural evolution, have received widespread acclaim for their exceptional performance in various real-world optimization problems. While…
Algorithms for reinforcement learning (RL) in large state spaces crucially rely on supervised learning subroutines to estimate objects such as value functions or transition probabilities. Since only the simplest supervised learning problems…
We revisit the classic Cournot model and extend it to a two-echelon supply chain with an upstream supplier who operates under demand uncertainty and multiple downstream retailers who compete over quantity. The supplier's belief about retail…
We establish new tail estimates for order statistics and for the Euclidean norms of projections of an isotropic log-concave random vector. More generally, we prove tail estimates for the norms of projections of sums of independent…
Pairwise Ranking Prompting (PRP) elicits pairwise preference judgments from an LLM, which are then aggregated into a ranking, usually via classical sorting algorithms. However, judgments are noisy, order-sensitive, and sometimes…