English

Efficient Reasoning via Thought-Training and Thought-Free Inference

Computation and Language 2025-12-01 v3

Abstract

Recent advances in large language models (LLMs) have leveraged explicit Chain-of-Thought (CoT) prompting to improve reasoning accuracy. However, most existing methods primarily focus on compressing verbose reasoning outputs. These Long-to-Short transformations aim to improve efficiency, but require a large amount of short CoT data. In this work, we introduce \textbf{3TF} (\textbf{T}hought-\textbf{T}raining and \textbf{T}hought-\textbf{F}ree inference), a framework for efficient reasoning that takes a Short-to-Long perspective. We first train a hybrid model that can operate in both reasoning and non-reasoning modes, and then further train it on CoT-annotated data to internalize structured reasoning, while enforcing concise, thought-free outputs at inference time using the no-reasoning mode. Unlike compression-based approaches, 3TF improves the reasoning quality of non-reasoning outputs, enabling models to perform rich internal reasoning implicitly while keeping external outputs short. Empirically, 3TF-trained models obtain large improvements on reasoning benchmarks under thought-free inference, demonstrating that high quality reasoning can be learned and executed implicitly without explicit step-by-step generation.

Keywords

Cite

@article{arxiv.2511.03408,
  title  = {Efficient Reasoning via Thought-Training and Thought-Free Inference},
  author = {Canhui Wu and Qiong Cao and Chao Xue and Wei Xi and Xiaodong He},
  journal= {arXiv preprint arXiv:2511.03408},
  year   = {2025}
}

Comments

11 pages, 4 figures

R2 v1 2026-07-01T07:22:45.633Z