Related papers: MRL order, log-concavity and an application to pea…
Learning to rank is a supervised learning problem where the output space is the space of rankings but the supervision space is the space of relevance scores. We make theoretical contributions to the learning to rank problem both in the…
In this paper, we consider the problem of order preservation under addition and multiplication operators over the vector space of univariate real-valued random variables. Consistent with the case of usual order over the real numbers-as…
We give a necessary and sufficient condition for a homogeneous Markov process taking values in $\R^n$ to enjoy the time-inversion property of degree $\alpha$. The condition sets the shape for the semigroup densities of the process and…
The time evolution of complex systems usually can be described through stochastic processes. These processes are measured at finite resolution, what necessarily reduces them to finite sequences of real numbers. In order to relate these data…
Process Reward Modeling (PRM) is critical for complex reasoning and decision-making tasks where the accuracy of intermediate steps significantly influences the overall outcome. Existing PRM approaches, primarily framed as classification…
Despite the remarkable progress of multimodal large language models (MLLMs), they continue to face challenges in achieving competitive performance on ordinal regression (OR; a.k.a. ordinal classification). To address this issue, this paper…
Training large language models with reinforcement learning (RL) against verifiable rewards significantly enhances their reasoning abilities, yet remains computationally expensive due to inefficient uniform prompt sampling. We introduce…
Retrieval plays a fundamental role in recommendation systems, search, and natural language processing (NLP) by efficiently finding relevant items from a large corpus given a query. Dot products have been widely used as the similarity…
Designing sample-efficient and computationally feasible reinforcement learning (RL) algorithms is particularly challenging in environments with large or infinite state and action spaces. In this paper, we advance this effort by presenting…
We show that PLDR-LLMs pretrained at self-organized criticality exhibit reasoning at inference time. The characteristics of PLDR-LLM deductive outputs at criticality is similar to second-order phase transitions. At criticality, the…
This paper studies the continuous-time q-learning (the continuous time counterpart of Q-learing) for Markov switching system under Tsallis entropy regularization. We address the difficulty in traditional RL algorithms where the Tsallis…
It is known that the Azema-Yor solution to the Skorokhod embedding problem maximizes the law of the running maximum of an uniformly integrable martingale with given terminal value distribution. Recently this optimality property has been…
In many data analysis pipelines, a basic and time-consuming process is to produce join results and feed them into downstream tasks. Numerous enumeration algorithms have been developed for this purpose. To be a statistically meaningful…
We consider continuous-time Markov chains which display a family of wells at the same depth. We provide sufficient conditions which entail the convergence of the finite-dimensional distributions of the order parameter to the ones of a…
Given a general critical or sub-critical branching mechanism, we define a pruning procedure of the associated L\'evy continuum random tree. This pruning procedure is defined by adding some marks on the tree, using L\'evy snake techniques.…
We introduce methodology for real-time inference in general-state-space hidden Markov models. Specifically, we extend recent advances in controlled sequential Monte Carlo (CSMC) methods-originally proposed for offline smoothing-to the…
We introduce an expressive subclass of non-negative almost submodular set functions, called strongly 2-coverage functions which include coverage and (sums of) matroid rank functions, and prove that the homogenization of the generating…
Although the foundations of ranking are well established, the ranking literature has primarily been focused on simple, unimodal models, e.g. the Mallows and Plackett-Luce models, that define distributions centered around a single total…
We introduce the notion of interlacing log-concavity of a polynomial sequence $\{P_m(x)\}_{m\geq 0}$, where $P_m(x)$ is a polynomial of degree m with positive coefficients $a_{i}(m)$. This sequence of polynomials is said to be interlacing…
We study ergodic properties of a class of Markov-modulated general birth-death processes under fast regime switching. The first set of results concerns the ergodic properties of the properly scaled joint Markov process with a parameter that…