DeepSeek:大型 AI 模型中的范式转变与技术演进
人工智能
2025-07-15 v1
摘要
DeepSeek 是一家中国人工智能公司,已发布其 V3 和 R1 系列模型,由于低成本、高性能和开源优势吸引了全球关注。本文首先回顾了大型 AI 模型的演化,聚焦于范式转变、主流的大型语言模型(LLM)范式以及 DeepSeek 的范式。随后,本文突出了 DeepSeek 引入的新型算法,包括 Multi-head Latent Attention(MLA)、Mixture-of-Experts(MoE)、Multi-Token Prediction(MTP)以及 Group Relative Policy Optimization(GRPO)。本文进一步探讨了 DeepSeek 在 LLM 规模化、训练、推理和系统级优化架构方面的工程突破。此外,分析了 DeepSeek 模型对竞争性 AI 生态系统的影响,将其与主流 LLM 在各领域进行比较。最后,本文反思了从 DeepSeek 创新中获得的洞见,并讨论了大型 AI 模型在数据、训练和推理方面的未来趋势。
关键词
引用
@article{arxiv.2507.09955,
title = {DeepSeek: Paradigm Shifts and Technical Evolution in Large AI Models},
author = {Luolin Xiong and Haofen Wang and Xi Chen and Lu Sheng and Yun Xiong and Jingping Liu and Yanghua Xiao and Huajun Chen and Qing-Long Han and Yang Tang},
journal= {arXiv preprint arXiv:2507.09955},
year = {2025}
}