中文

面向众核处理器上深度学习的性能调优(硕士论文)

分布式、并行与集群计算 2018-06-05 v1

摘要

卷积神经网络(CNNs)正因在各种应用中的成功而变得非常流行。Loki 众核处理器架构在实现专用硬件性能与效率的同时作为一种通用解决方案极具前景。Loki 将许多简单核心与增强的程序员控制相结合。这种自由度可被用来产生比传统多处理器高效得多的代码,但也创造了一个用于可能优化的非常庞大的设计空间。在本项目中,我探索了 CNN 应用的可能优化、它们在不同 Loki 特定配置、卷积参数和输入上的可移植性。最后,我研究了用于进一步提升性能的自适应算法的潜力。

关键词

引用

@article{arxiv.1806.01105,
  title  = {Performance tuning for deep learning on a many-core processor (master thesis)},
  author = {Philippos Papaphilippou},
  journal= {arXiv preprint arXiv:1806.01105},
  year   = {2018}
}

备注

A dissertation submitted to the University of Cambridge on June 16, 2017 in partial fulfilment of the requirements for the degree of Master of Philosophy in Advanced Computer Science