中文

超叠加梯度下降: Harnessing 量子原理进行模型训练

机器学习 2026-01-30 v2 量子物理

摘要

大型语言模型(LLM)日益采用类似 AdamW 的经典优化技术以提升收敛速度与泛化能力。然而,量子启发方法如何增强经典训练的机制仍未得到充分探讨。我们提出超叠加梯度下降(Superpositional Gradient Descent, SGD),这是一种新型优化器,通过注入量子电路扰动将梯度更新与量子叠加关联。我们构建了数学框架,并在 PyTorch 与 Qiskit 中实现了混合量子-经典电路。于合成序列分类和大规模 LLM 微调任务上,SGD 收敛速度更快且最终损失更低于 AdamW。尽管结果令人鼓舞,但规模化与硬件限制仍制约其实际应用。总体而言,本工作为量子计算与深度学习交叉领域提供了新见解,指明了利用量子原理控制与提升模型行为的可行路径。

关键词

引用

@article{arxiv.2511.01918,
  title  = {Superpositional Gradient Descent: Harnessing Quantum Principles for Model Training},
  author = {Ahmet Erdem Pamuk and Emir Kaan Özdemir and Şuayp Talha Kocabay},
  journal= {arXiv preprint arXiv:2511.01918},
  year   = {2026}
}

备注

Accepted at 2025 IEEE International Conference on Quantum Artificial Intelligence (IEEE QAI 2025). This is the accepted version of the paper. The final published version will appear in the IEEE proceedings. \c{opyright} 2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses