English

Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors

Artificial Intelligence 2025-10-28 v1

Abstract

We present a novel approach for controllable mathematical reasoning that leverages self-optimizing thought vectors with entropy minimization. Our method introduces learnable thought vectors that dynamically modulate the internal reasoning process of large language models. Using Gemma-2-9B on GSM8K, we achieve 90.1% accuracy with a controllability score of 0.42, demonstrating that entropy-based rewards effectively guide focused reasoning patterns without requiring external reward annotations. Our analysis reveals distinct thought vector clusters and consistent low-entropy distributions across control conditions, validating our framework for controllable AI reasoning.

Keywords

Cite

@article{arxiv.2510.22132,
  title  = {Controllable Mathematical Reasoning via Self-Optimizing Thought Vectors},
  author = {Xuying LI},
  journal= {arXiv preprint arXiv:2510.22132},
  year   = {2025}
}
R2 v1 2026-07-01T07:05:13.661Z