中文
相关论文

相关论文: Rethinking GSPO: The Perplexity-Entropy Equivalenc…

200 篇论文

We consider discrete stochastic processes, modeled by classical master equations, on networks. The temporal growth of the lack of information about the system is captured by its non-equilibrium entropy, defined via the transition…

统计力学 · 物理学 2017-04-26 Oliver Muelken , Sarah Heinzelmann , Maxim Dolgushev

Sequential recommender systems have achieved steady gains in offline accuracy, yet it remains unclear how close current models are to the intrinsic accuracy limit imposed by the data. A reliable, model-agnostic estimate of this ceiling…

信息检索 · 计算机科学 2026-04-15 En Xu , Jingtao Ding , Yong Li

Motivation: Entropy measurements on hierarchical structures have been used in methods for information retrieval and natural language modeling. Here we explore its application to semantic similarity. By finding shared ontology terms,…

计算与语言 · 计算机科学 2017-06-20 Andrew Warren , Joao Setubal

Direct Preference Optimization (DPO) has been successfully used to align large language models (LLMs) according to human preferences, and more recently it has also been applied to improving the quality of text-to-image diffusion models.…

计算机视觉与模式识别 · 计算机科学 2025-10-07 Shivanshu Shekhar , Shreyas Singh , Tong Zhang

Many reinforcement learning (RL) problems admit multiple terminal solutions of comparable quality, where the goal is not to identify a single optimum but to represent a diverse set of high-quality outcomes. Nevertheless, policies trained by…

机器学习 · 计算机科学 2026-01-30 Abhijeet Sinha , Sundari Elango , Dianbo Liu

Existing alignment paradigms remain limited in capturing the pluralistic nature of human values. Overton Pluralism addresses this gap by generating responses with diverse perspectives from a single query. This paper introduces OP-GRPO…

计算与语言 · 计算机科学 2026-02-25 Yu Fu , Seongho Son , Ilija Bogunovic

Importance weighting is a classic technique to handle distribution shifts. However, prior work has presented strong empirical and theoretical evidence demonstrating that importance weights can have little to no effect on overparameterized…

机器学习 · 计算机科学 2022-03-07 Ke Alexander Wang , Niladri S. Chatterji , Saminul Haque , Tatsunori Hashimoto

Recently, information theoretic analysis has become a popular framework for understanding the generalization behavior of deep neural networks. It allows a direct analysis for stochastic gradient/Langevin descent (SGD/SGLD) learning…

机器学习 · 统计学 2023-05-03 Yuxin Dong , Tieliang Gong , Hong Chen , Chen Li

The Random Permutation Set (RPS) is a new type of set proposed recently, which can be regarded as the generalization of evidence theory. To measure the uncertainty of RPS, the entropy of RPS and its corresponding maximum entropy have been…

信息论 · 计算机科学 2024-03-12 Jiefeng Zhou , Zhen Li , Kang Hao Cheong , Yong Deng

Recent literature in the last Maximum Entropy workshop introduced an analogy between cumulative probability distributions and normalized utility functions. Based on this analogy, a utility density function can de defined as the derivative…

人工智能 · 计算机科学 2009-11-10 Ali E. Abbas

Entropy always increases monotonically in a closed system but complexity increases at first and then decreases as equilibrium is approached. Commonsense information-related definitions for entropy and complexity demonstrate that complexity…

物理与社会 · 物理学 2025-03-26 Theodore Modis

Reinforcement learning with verifiable rewards (RLVR) has become a practical route to improve large language model reasoning, and Group Relative Policy Optimization (GRPO) is a widely used optimizer in this setting. However, RLVR training…

机器学习 · 计算机科学 2026-05-14 Tue Le , Linh Ngo Van , Trung Le

Group Relative Policy Optimization (GRPO) has become the dominant method for reinforcement learning with verifiable rewards in large language models, but it suffers from two critical limitations: gradient vanishing and diversity collapse.…

机器学习 · 计算机科学 2026-05-20 Khiem Le , Phuc Nguyen , Youssef Mroueh , Chi-Heng Lin , Shangqian Gao , Ting Hua , Nitesh V. Chawla

Learned sparse retrieval models such as SPLADE combine the effectiveness of neural architectures with the efficiency of inverted indices. As these models assign weights to terms from a fixed vocabulary, interpretability is often touted as a…

信息检索 · 计算机科学 2026-05-20 Gregory Polyakov , Harrisen Scells , Carsten Eickhoff

We introduce the problem of \emph{entropy equivalence testing} for probability distributions, a relaxation of the well-studied closeness testing problem, where the distribution testing algorithm is now only required to distinguish, given…

数据结构与算法 · 计算机科学 2026-05-25 Clément L. Canonne , Yash Pote , Jonathan Scarlett , Joy Qiping Yang

Nearly all practical neural models for classification are trained using cross-entropy loss. Yet this ubiquitous choice is supported by little theoretical or empirical evidence. Recent work (Hui & Belkin, 2020) suggests that training using…

机器学习 · 计算机科学 2023-02-09 Like Hui , Mikhail Belkin , Stephen Wright

Entropy is a measure of heterogeneity widely used in applied sciences, often when data are collected over space. Recently, a number of approaches has been proposed to include spatial information in entropy. The aim of entropy is to…

统计理论 · 数学 2019-11-12 Linda Altieri , Daniela Cocchi , Giulia Roli

We consider the weighted least squares spline approximation of a noisy dataset. By interpreting the weights as a probability distribution, we maximize the associated entropy subject to the constraint that the mean squared error is…

数值分析 · 数学 2024-01-19 Luigi Brugnano , Domenico Giordano , Felice Iavernaro , Giorgia Rubino

Post-training has become central to improving reasoning and alignment in large language models, where critic-free models enable scalable learning from model-generated outputs but lack principled mechanisms to distinguish informative from…

机器学习 · 计算机科学 2026-05-22 Zheyuan Zhang , Kaiwen Shi , Han Bao , Zehong Wang , Tianyi Ma , Yanfang Ye

When the available information is noisy zeroth-order (ZO) oracle, stochastic approximation methods are popular for estimating the root of the multivariate gradient equation. Inspired by the Stein's identity, this work establishes a novel…

最优化与控制 · 数学 2021-04-06 Jingyi Zhu