中文
相关论文

相关论文: Rethinking GSPO: The Perplexity-Entropy Equivalenc…

200 篇论文

Recently, substantial research efforts in Deep Metric Learning (DML) focused on designing complex pairwise-distance losses, which require convoluted schemes to ease optimization, such as sample mining or pair weighting. The standard…

The concept of Entropy plays a key role in Information Theory, Statistics, and Machine Learning.This paper introduces a new entropy measure, called the t-entropy, which exploits the concavity of the inverse-tan function. We analytically…

信息论 · 计算机科学 2021-05-06 Saptarshi Chakraborty , Debolina Paul , Swagatam Das

Rooted trees with probabilities are used to analyze properties of a variable length code. A bound is derived on the difference between the entropy rates of the code and a memoryless source. The bound is in terms of normalized informational…

信息论 · 计算机科学 2013-10-11 Georg Böcherer , Rana Ali Amjad

Post-training LLMs with Reinforcement Learning, specifically Group Relative Policy Optimization (GRPO), has emerged as a paradigm for enhancing mathematical reasoning. However, standard GRPO relies on scalar correctness rewards that are…

计算与语言 · 计算机科学 2026-03-03 Xiwen Chen , Wenhui Zhu , Peijie Qiu , Xuanzhao Dong , Hao Wang , Haiyu Wu , Huayu Li , Aristeidis Sotiras , Yalin Wang , Abolfazl Razi

Proximal Policy Optimization (PPO) is among the most widely used deep reinforcement learning algorithms, yet its theoretical foundations remain incomplete. Most importantly, convergence and understanding of fundamental PPO advantages remain…

We extend previously proposed measures of complexity, emergence, and self-organization to continuous distributions using differential entropy. This allows us to calculate the complexity of phenomena for which distributions are known. We…

适应与自组织系统 · 物理学 2016-04-01 Guillermo Santamaría-Bonfil , Nelson Fernández , Carlos Gershenson

RLVR has become a widely adopted paradigm for improving LLMs' reasoning capabilities, and GRPO is one of its most representative algorithms. In this paper, we first show that GRPO admits an equivalent discriminative reformulation as a…

机器学习 · 计算机科学 2026-05-19 Feng Zhang , Xinhong Ma , Ziqiang Dong , Xi Leng , Jianfei Zhao , Xin Sun , Yang Yang , Guanjun Jiang

The diversity of the symbols of the information source is calculated following the definition that entropy is the information loss and following a new entropy-symbol similarity relation after the rejection of the Gibbs paradox statement.…

数据分析、统计与概率 · 物理学 2007-05-23 Shu-Kun Lin

It is not obvious how to extend Shannon's original information entropy to higher dimensions, and many different approaches have been tried. We replace the English text symbol sequence originally used to illustrate the theory by a discrete,…

信息论 · 计算机科学 2016-09-06 Kieran G. Larkin

In computer vision, it is often observed that formulating regression problems as a classification task often yields better performance. We investigate this curious phenomenon and provide a derivation to show that classification, with the…

计算机视觉与模式识别 · 计算机科学 2023-03-01 Shihao Zhang , Linlin Yang , Michael Bi Mi , Xiaoxu Zheng , Angela Yao

Large Language Models (LLMs) have made remarkable progress in enhancing step-by-step reasoning through reinforcement learning. However, the Group Relative Policy Optimization (GRPO) algorithm, which relies on sparse reward rules, often…

人工智能 · 计算机科学 2025-07-30 Xingjian Zhang , Siwei Wen , Wenjun Wu , Lei Huang

Permutation entropy quantifies the diversity of possible orderings of the values a random or deterministic system can take, as Shannon entropy quantifies the diversity of values. We show that the metric and permutation entropy…

混沌动力学 · 物理学 2016-08-16 Jose M. Amigo , Matthew B. Kennel , Ljupco Kocarev

Can we learn more from data than existed in the generating process itself? Can new and useful information be constructed from merely applying deterministic transformations to existing data? Can the learnable content in data be evaluated…

机器学习 · 计算机科学 2026-03-17 Marc Finzi , Shikai Qiu , Yiding Jiang , Pavel Izmailov , J. Zico Kolter , Andrew Gordon Wilson

Recently, GRPO-based reinforcement learning has shown remarkable progress in optimizing flow-matching models, effectively improving their alignment with task-specific rewards. Within these frameworks, the policy update relies on…

计算机视觉与模式识别 · 计算机科学 2025-10-31 Jing Wang , Jiajun Liang , Jie Liu , Henglin Liu , Gongye Liu , Jun Zheng , Wanyuan Pang , Ao Ma , Zhenyu Xie , Xintao Wang , Meng Wang , Pengfei Wan , Xiaodan Liang

In previous work (see arxiv:1102.3040), we have defined the telescopic relative entropy (TRE), which is a regularisation of the quantum relative entropy $S(\rho||\sigma)=\trace\rho(\log\rho-\log\sigma)$, by replacing the second argument…

数学物理 · 物理学 2011-04-28 Koenraad M. R. Audenaert

The ability to quantify the directional flow of information is vital to understanding natural systems and designing engineered information-processing systems. A widely used measure to quantify this information flow is the transfer entropy.…

分子网络 · 定量生物学 2025-07-11 Avishek Das , Pieter Rein ten Wolde

Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs), particularly by boosting their mathematical performance. However, GRPO and related entropy-regularization methods…

计算与语言 · 计算机科学 2025-10-13 Xingyu Lin , Yilin Wen , En Wang , Du Su , Wenbin Liu , Chenfu Bao , Zhonghou Lv

The extropy is a measure of information introduced by Lad et al. (2015) as dual to entropy. As the entropy, it is a shift-independent information measure. We introduce here the notion of weighted extropy, a shift-dependent information…

统计理论 · 数学 2020-08-19 Narayanaswamy Balakrishnan , Francesco Buono , Maria Longobardi

Entropy is a measure of self-information which is used to quantify losses. Entropy was developed in thermodynamics, but is also used to compare probabilities based on their deviating information content. Corresponding model uncertainty is…

概率论 · 数学 2018-01-23 Alois Pichler , Ruben Schlotter

Similarity-sensitive entropy measures the uncertainty of a probability law relative to a similarity kernel that encodes the distinguishability between states. We develop a measure-theoretic treatment covering both finite similarity matrices…

概率论 · 数学 2026-05-29 Joseph Samuel Miller