中文

Falcon-H1R:通过混合模型实现推理前沿的高效测试时扩展

人工智能 2026-01-06 v1

摘要

本 work 引入 Falcon-H1R,一个 7B 参数的 reasoning-optimized 模型, establish 了以小型语言模型 (SLM) 达到 competitive reasoning performance 的可行性。Falcon-H1R 在参数效率方面表现突出, across 多种 reasoning-intensive benchmark,始终匹配或超过 2倍到 7倍 更大模型的 SOTA 推理模型。这些结果 underscore 在不增加模型规模的情况下,通过 careful data curation 和 targeted training strategy (通过 efficient SFT 和 RL scaling) 能实现显著性能提升的重要性。此外,Falcon-H1R 通过结合 faster inference (通过其 hybrid-parallel architecture design)、token efficiency 和 higher accuracy, advance 3D reasoning efficiency 的 limits。这一 unique blend 使 Falcon-H1R-7B 成为 scaling advanced reasoning systems 的 practical backbone,特别是在需要 extensive chain-of-thoughts generation 和 parallel test-time scaling 的场景中。凭借最近引入的 DeepConf 方法,Falcon-H1R 实现了 state-of-the-art test-time scaling efficiency, 在 accuracy 和 computational cost 方面均实现显著改进。结果表明,compact models 通过 targeted model training 和 architectural choices 能交付 robust and scalable reasoning performance。

关键词

引用

@article{arxiv.2601.02346,
  title  = {Falcon-H1R: Pushing the Reasoning Frontiers with a Hybrid Model for Efficient Test-Time Scaling},
  author = {Falcon LLM Team and Iheb Chaabane and Puneesh Khanna and Suhail Mohmad and Slim Frikha and Shi Hu and Abdalgader Abubaker and Reda Alami and Mikhail Lubinets and Mohamed El Amine Seddik and Hakim Hacid},
  journal= {arXiv preprint arXiv:2601.02346},
  year   = {2026}
}