中文

大语言模型从概念到实现综述

计算与语言 2024-05-29 v2 人工智能 信息论 机器学习 math.IT

摘要

大语言模型 (LLM) 的最新进展,特别是基于 Transformer 架构构建的模型,显著拓宽了自然语言处理 (NLP) 应用的范围,超越了其在聊天机器人技术中的初始用途。本文研究了这些模型的多方面应用,重点强调了 GPT 系列。这一探索重点关注人工智能 (AI) 驱动的工具在变革编码和问题求解等传统任务方面的变革性影响,同时也为跨行业的研究与开发开辟了新路径。从代码解释和图像描述到促进交互式系统的构建和推进计算领域,Transformer 模型体现了深度学习、数据分析和神经网络设计的协同。本综述深入探讨了 Transformer 模型的最新研究,突出了其多功能性及其在变革多样化应用领域的潜力,从而为读者提供了基于 Transformer 的 LLM 在实际应用中的当前和未来全景的全面理解。

关键词

引用

@article{arxiv.2403.18969,
  title  = {A Survey on Large Language Models from Concept to Implementation},
  author = {Chen Wang and Jin Zhao and Jiaqi Gong},
  journal= {arXiv preprint arXiv:2403.18969},
  year   = {2024}
}

备注

Section 3 lacks to clarity and accuracy in defining the applications and capabilities of LLMs. More rework needs to be done on illustrate how LLMs being used in cross-domains