Asymmetric processors have emerged as an appealing technology for severely energy-constrained environments, especially in the mobile market where heterogeneity in applications is mainstream. In addition, given the growing interest on ultra low-power architectures for high performance computing, this type of platforms are also being investigated in the road towards the implementation of energy- efficient high-performance scientific applications. In this paper, we propose a first step towards a complete implementation of the BLAS interface adapted to asymmetric ARM big.LITTLE processors, analyzing the trade-offs between performance and energy efficiency when compared to existing homogeneous (symmetric) multi-threaded BLAS implementations. Our experimental results reveal important gains in performance while maintaining the energy efficiency of homogeneous solutions by efficiently exploiting all the resources of the asymmetric processor.
@article{arxiv.1507.05129,
title = {Performance and Energy Optimization of Matrix Multiplication on Asymmetric big.LITTLE Processors},
author = {Sandra Catalán and Francisco D. Igual and Rafael Mayo and Luis Piñuel and Enrique S. Quintana-Ortí and Rafael Rodríguez-Sánchez},
journal= {arXiv preprint arXiv:1507.05129},
year = {2015}
}
Comments
Presented at HiPEAC 2015, Amsterdam. Foundation of the Asymmetric BLIS implementation