什么跟在 transformer 之后?——连接深度学习想法的选择性综述
摘要
自 2017 年以来,transformer 已成为人工智能领域事实上的标准模型,尽管存在从能源效率到幻觉等诸多局限。研究在改进 transformer 的各个方面取得了很多进展,更广泛地涉及深度学习,涌现出许多关于架构、层、优化目标和优化技术的提议。对于研究者来说,跟踪这些发展在更广泛层面上具有挑战性。我们提供了对近期重要工作的全面概述,供已具备深度学习基本理解的读者使用。我们的关注点不同于其他工作,特别针对新兴的、可能颠覆性的 transformer 替代方案以及近期深度学习的成功想法。我们希望,这种对有影响力的近期工作和新想法的整合、统一的处理能帮助研究者在深度学习的不同领域之间形成新的联系。我们识别并讨论了多个模式,概括过去十年成功创新的关键策略,以及可以视为新星的作品。特别是,我们讨论了如何改进 transformer 的尝试,涵盖(部分)已证实的方法,如状态空间模型,也包括尽管未实现最先进结果却似乎有前景的深远想法。我们还涵盖了关于近期最先进模型的讨论,如 OpenAI 的 GPT 系列、Meta 的 LLaMA 模型以及 Google 的 Gemini 模型家族。
关键词
引用
@article{arxiv.2408.00386,
title = {What comes after transformers? -- A selective survey connecting ideas in deep learning},
author = {Johannes Schneider},
journal= {arXiv preprint arXiv:2408.00386},
year = {2024}
}
备注
This is an extended version of the published paper by Johannes Schneider and Michalis Vlachos titled "A survey of deep learning: From activations to transformers'' which appeared at the International Conference on Agents and Artificial Intelligence(ICAART) in 2024. It was selected for post-publication and has been submitted to the post-publication proceedings