Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space
Abstract
Large Language Models (LLMs) apply uniform computation to all tokens, despite language exhibiting highly non-uniform information density. This token-uniform regime wastes capacity on locally predictable spans while under-allocating computation to semantically critical transitions. We propose , a hierarchical language modeling framework that learns semantic boundaries from latent representations and shifts computation from tokens to a compressed concept space where reasoning is more efficient. DLCM discovers variable-length concepts end-to-end without relying on predefined linguistic units. Hierarchical compression fundamentally changes scaling behavior. We introduce the first , which disentangles token-level capacity, concept-level reasoning capacity, and compression ratio, enabling principled compute allocation under fixed FLOPs. To stably train this heterogeneous architecture, we further develop a \textbf{decoupled \muP parametrization} that supports zero-shot hyperparameter transfer across widths and compression regimes. At a practical setting (, corresponding to an average of four tokens per concept), DLCM reallocates roughly one-third of inference compute into a higher-capacity reasoning backbone, achieving a \textbf{+2.69\% average improvement} across 12 zero-shot benchmarks under matched inference FLOPs.
Cite
@article{arxiv.2512.24617,
title = {Dynamic Large Concept Models: Latent Reasoning in an Adaptive Semantic Space},
author = {Xingwei Qu and Shaowen Wang and Zihao Huang and Kai Hua and Fan Yin and Rui-Jie Zhu and Jundong Zhou and Qiyang Min and Zihao Wang and Yizhi Li and Tianyu Zhang and He Xing and Zheng Zhang and Yuxuan Song and Tianyu Zheng and Zhiyuan Zeng and Chenghua Lin and Ge Zhang and Wenhao Huang},
journal= {arXiv preprint arXiv:2512.24617},
year = {2026}
}