English

On the Transformations across Reward Model, Parameter Update, and In-Context Prompt

Computation and Language 2024-06-25 v1 Artificial Intelligence

Abstract

Despite the general capabilities of pre-trained large language models (LLMs), they still need further adaptation to better serve practical applications. In this paper, we demonstrate the interchangeability of three popular and distinct adaptation tools: parameter updating, reward modeling, and in-context prompting. This interchangeability establishes a triangular framework with six transformation directions, each of which facilitates a variety of applications. Our work offers a holistic view that unifies numerous existing studies and suggests potential research directions. We envision our work as a useful roadmap for future research on LLMs.

Keywords

Cite

@article{arxiv.2406.16377,
  title  = {On the Transformations across Reward Model, Parameter Update, and In-Context Prompt},
  author = {Deng Cai and Huayang Li and Tingchen Fu and Siheng Li and Weiwen Xu and Shuaiyi Li and Bowen Cao and Zhisong Zhang and Xinting Huang and Leyang Cui and Yan Wang and Lemao Liu and Taro Watanabe and Shuming Shi},
  journal= {arXiv preprint arXiv:2406.16377},
  year   = {2024}
}
R2 v1 2026-06-28T17:16:52.198Z