English

PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching

Computation and Language 2026-05-26 v2 Artificial Intelligence Machine Learning

Abstract

Unsupervised Reinforcement Learning from Internal Feedback (RLIF) has emerged as a promising paradigm for eliciting the latent capabilities of Large Language Models (LLMs) without external supervision. However, current methods rely on heuristic intrinsic rewards, which often lack a well-defined theoretical optimization target and are prone to degenerative biases. In this work, we introduce PowerFlow, a principled framework that reformulates unsupervised fine-tuning as a distribution matching problem. By casting GFlowNet as an amortized variational sampler for unnormalized densities, we propose a length-aware Trajectory-Balance objective that explicitly neutralizes the structural length biases inherent in autoregressive generation. By targeting α\alpha-power distributions, PowerFlow enables the directional elicitation of the dual nature of LLMs: sharpening the distribution (α>1\alpha > 1) to intensify logical reasoning, or flattening it (α<1\alpha < 1) to unlock expressive creativity. Extensive experiments demonstrate that PowerFlow consistently outperforms existing RLIF methods, matching or even exceeding supervised GRPO. Furthermore, by mitigating over-sharpening in aligned models, our approach achieves simultaneous gains in diversity and quality, shifting the Pareto frontier in creative tasks.

Keywords

Cite

@article{arxiv.2603.18363,
  title  = {PowerFlow: Unlocking the Dual Nature of LLMs via Principled Distribution Matching},
  author = {Ruishuo Chen and Yu Chen and Zhuoran Li and Longbo Huang},
  journal= {arXiv preprint arXiv:2603.18363},
  year   = {2026}
}

Comments

Camera-ready version accepted at ICML 2026