中文

TeleChat2、TeleChat2.5与T1技术报告

计算与语言 2025-07-30 v3

摘要

我们介绍最新的一系列TeleChat模型:\textbf{TeleChat2}、\textbf{TeleChat2.5}和\textbf{T1},相较于其前身TeleChat实现了显著的性能提升。尽管对模型架构的更改较小,但通过在预训练和后训练阶段的增强训练策略,实现了显著的性能提升。该系列从\textbf{TeleChat2}开始,在10万亿个高质量且多样化的标记上进行预训练。随后进行监督微调(SFT)和直接偏好优化(DPO)以进一步提升其能力。\textbf{TeleChat2.5}和\textbf{T1}通过纳入以特定领域数据集为中心的持续预训练阶段,并结合强化学习(RL)来提高代码生成和数学推理任务的性能。 designed for complex reasoning, supporting long Chain-of-Thought (CoT) reasoning and demonstrating substantial improvements in mathematics and coding. In contrast, \textbf{TeleChat2.5} prioritizes speed, delivering rapid inference. Both flagship models of \textbf{T1} and \textbf{TeleChat2.5} are dense Transformer-based architectures with 115B parameters, showcasing significant advancements in reasoning and general task performance compared to the original TeleChat. Notably, \textbf{T1-115B} outperform proprietary models such as OpenAI's o1-mini and GPT-4o. We publicly release \textbf{TeleChat2}, \textbf{TeleChat2.5} and \textbf{T1}, including post-trained versions with 35B and 115B parameters, to empower developers and researchers with state-of-the-art language models tailored for diverse applications.

关键词

引用

@article{arxiv.2507.18013,
  title  = {Technical Report of TeleChat2, TeleChat2.5 and T1},
  author = {Zihan Wang and Xinzhang Liu and Yitong Yao and Chao Wang and Yu Zhao and Zhihao Yang and Wenmin Deng and Kaipeng Jia and Jiaxin Peng and Yuyao Huang and Sishi Xiong and Zhuo Jiang and Kaidong Yu and Xiaohui Hu and Fubei Yao and Ruiyu Fang and Zhuoru Jiang and Ruiting Song and Qiyi Xie and Rui Xue and Xuewei He and Yanlei Xue and Zhu Yuan and Zhaoxi Zhang and Zilu Huang and Shiquan Wang and Xin Wang and Hanming Wu and Mingyuan Wang and Xufeng Zhan and Yuhan Sun and Zhaohu Xing and Yuhao Jiang and Bingkai Yang and Shuangyong Song and Yongxiang Li and Zhongjiang He and Xuelong Li},
  journal= {arXiv preprint arXiv:2507.18013},
  year   = {2025}
}

备注

32 pages, 5 figures