English

AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models

Artificial Intelligence 2025-10-13 v4 Machine Learning

Abstract

LLMs often need effective configurations, like temperature and reasoning steps, to handle tasks requiring sophisticated reasoning and problem-solving, ranging from joke generation to mathematical reasoning. Existing prompting approaches usually adopt general-purpose, fixed configurations that work 'well enough' across tasks but seldom achieve task-specific optimality. To address this gap, we introduce AdaReasoner, an LLM-agnostic plugin designed for any LLM to automate adaptive reasoning configurations for tasks requiring different types of thinking. AdaReasoner is trained using a reinforcement learning (RL) framework, combining a factorized action space with a targeted exploration strategy, along with a pretrained reward model to optimize the policy model for reasoning configurations with only a few-shot guide. AdaReasoner is backed by theoretical guarantees and experiments of fast convergence and a sublinear policy gap. Across six different LLMs and a variety of reasoning tasks, it consistently outperforms standard baselines, preserves out-of-distribution robustness, and yield gains on knowledge-intensive tasks through tailored prompts.

Keywords

Cite

@article{arxiv.2505.17312,
  title  = {AdaReasoner: Adaptive Reasoning Enables More Flexible Thinking in Large Language Models},
  author = {Xiangqi Wang and Yue Huang and Yanbo Wang and Xiaonan Luo and Kehan Guo and Yujun Zhou and Xiangliang Zhang},
  journal= {arXiv preprint arXiv:2505.17312},
  year   = {2025}
}
R2 v1 2026-07-01T02:32:50.633Z