中文

DeepSeek:大型 AI 模型中的范式转变与技术演进

人工智能 2025-07-15 v1

摘要

DeepSeek 是一家中国人工智能公司,已发布其 V3 和 R1 系列模型,由于低成本、高性能和开源优势吸引了全球关注。本文首先回顾了大型 AI 模型的演化,聚焦于范式转变、主流的大型语言模型(LLM)范式以及 DeepSeek 的范式。随后,本文突出了 DeepSeek 引入的新型算法,包括 Multi-head Latent Attention(MLA)、Mixture-of-Experts(MoE)、Multi-Token Prediction(MTP)以及 Group Relative Policy Optimization(GRPO)。本文进一步探讨了 DeepSeek 在 LLM 规模化、训练、推理和系统级优化架构方面的工程突破。此外,分析了 DeepSeek 模型对竞争性 AI 生态系统的影响,将其与主流 LLM 在各领域进行比较。最后,本文反思了从 DeepSeek 创新中获得的洞见,并讨论了大型 AI 模型在数据、训练和推理方面的未来趋势。

关键词

引用

@article{arxiv.2507.09955,
  title  = {DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models},
  author = {Luolin Xiong and Haofen Wang and Xi Chen and Lu Sheng and Yun Xiong and Jingping Liu and Yanghua Xiao and Huajun Chen and Qing-Long Han and Yang Tang},
  journal= {arXiv preprint arXiv:2507.09955},
  year   = {2025}
}