Related papers: Rethinking GSPO: The Perplexity-Entropy Equivalenc…
Given an event log as a collection of recorded real-world process traces, process mining aims to automatically construct a process model that is both simple and provides a useful explanation of the traces. Conformance checking techniques…
Our capacity to process information depends on the computational power at our disposal. Information theory captures our ability to distinguish states or communicate messages when it is unconstrained with unrivaled beauty and elegance. For…
We propose a general approach to construct weighted likelihood estimating equations with the aim of obtain robust estimates. The weight, attached to each score contribution, is evaluated by comparing the statistical data depth at the model…
We carry out a numerical study of the bi-partite entanglement entropy in the gapped regime of two paradigmatic quantum spin chain models: the Ising chain in an external magnetic field and the anti-ferromagnetic XXZ model. The universal…
Quantum entropy and skew information play important roles in quantum information science. They are defined by the trace of the positive operators so that the trace inequalities often have important roles to develop the mathematical theory…
Importance weighting is a general way to adjust Monte Carlo integration to account for draws from the wrong distribution, but the resulting estimate can be highly variable when the importance ratios have a heavy right tail. This routinely…
In this article, we propose two classes of relative information measures based on extropy, viz., the generalized extropy similarity ratio (GESR) and generalized extropy divergence ratio (GEDR), that measure the similarity and discrepancy…
We consider a probability distribution depending on a real parameter $x$. As functions of $x$, the R\'enyi entropy and the Tsallis entropy can be expressed in terms of the associated index of coincidence $S(x)$. We establish recurrence…
We introduce an ambidextrous view of stochastic dynamical systems, comparing their forward-time and reverse-time representations and then integrating them into a single time-symmetric representation. The perspective is useful theoretically,…
We introduce a family of scale-invariant entropy statistics derived from logarithmically aggregated distance distributions of point processes, with prime numbers serving as a motivating example. The construction associates to each finite…
Wasserstein Policy Optimization (WPO) is a recently proposed reinforcement learning algorithm that leverages Wasserstein gradient flows to optimize stochastic policies in continuous action spaces. Despite its empirical success, the…
Reinforcement learning with verifiable rewards has shown notable effectiveness in enhancing large language models (LLMs) reasoning performance, especially in mathematics tasks. However, such improvements often come with reduced outcome…
The entropy of a graph is an information-theoretic quantity which expresses the complexity of a graph \cite{DM1,M}. After Shannon introduced the definition of entropy to information and communication, many generalizations of the entropy…
A unified combinatorial definition of the information content and entropy of different types of patterns, compatible with the traditional concepts of information and entropy, going beyond the limitations of Shannon information interpretable…
Reinforcement learning (RL) plays an increasingly important role in enhancing the reasoning capabilities of large language models (LLMs), yet stable and performant policy optimization remains challenging. Token-level importance ratios often…
Large language models frequently exhibit suboptimal performance on low resource languages, primarily due to inefficient subword segmentation and systemic training data imbalances. In this paper, we propose Variable Entropy Policy…
Some essential conceptual aspects that will fill some logical gaps of the frame to interpret the gravity as an entropic force was investigated, we focus on some crucial issues that didn't emphasized in Verlinde's original…
Information theoretic quantities play a central role in machine learning. The recent surge in the complexity of data and models has increased the demand for accurate estimation of these quantities. However, as the dimension grows the…
Reinforcement Learning with Verifiable Rewards (RLVR) has emerged as an indispensable paradigm for enhancing reasoning in Large Language Models (LLMs). However, standard policy optimization methods, such as Group Relative Policy…
Recent alignment work on Large Language Models (LLMs) suggests preference optimization can improve reasoning by shifting probability mass toward better solutions. We test this claim in a resource-constrained setting by applying GRPO with…