Hardware Architecture · Computer Science
Towards a Multi-array Architecture for Accelerating Large-scale Matrix Multiplication on FPGAs
Junzhong Shen, Yuran Qiao, You Huang, Mei Wen +1
2018-03-13
Distributed, Parallel, and Cluster Computing · Computer Science
Throughput Optimizations for FPGA-based Deep Neural Network Inference
Thorbjörn Posewsky, Daniel Ziener
2018-10-02
Machine Learning · Computer Science
Learning Modular Exponentiation with Transformers
David Demitri Africa, Sara M. Kapoor, Theo Simon Sorg, Challenger Mishra
2025-10-24
Distributed, Parallel, and Cluster Computing · Computer Science
Heterogeneous Highly Parallel Implementation of Matrix Exponentiation Using GPU
Chittampally Vasanth Raja, Srinivas Balasubramanian, Prakash S Raghavendra
2012-04-16
Hardware Architecture · Computer Science
A Survey of FPGA-Based Neural Network Accelerator
Kaiyuan Guo, Shulin Zeng, Jincheng Yu, Yu Wang +1
2018-12-07
Machine Learning · Computer Science
Efficient Recurrent Neural Networks using Structured Matrices in FPGAs
Zhe Li, Shuo Wang, Caiwen Ding, Qinru Qiu +2
2018-03-26
Hardware Architecture · Computer Science
Fast, Scalable, Energy-Efficient Non-element-wise Matrix Multiplication on FPGA
Xuqi Zhu, Huaizhi Zhang, JunKyu Lee, Jiacheng Zhu +4
2024-07-09
Mathematical Physics · Physics
Evaluation of the matrix exponential function using finite elements in time
D H Gebremedhin, C A Weatherford, X Zhang, A Wynn +1
2008-11-18
Distributed, Parallel, and Cluster Computing · Computer Science
Accelerating Matrix Multiplication: A Performance Comparison Between Multi-Core CPU and GPU
Mufakir Qamar Ansari, Mudabir Qamar Ansari
2025-07-30
Machine Learning · Computer Science
FPGA-based Acceleration for Convolutional Neural Networks: A Comprehensive Review
Junye Jiang, Yaan Zhou, Yuanhao Gong, Haoxuan Yuan +1
2025-05-21
Distributed, Parallel, and Cluster Computing · Computer Science
Enabling OpenMP Task Parallelism on Multi-FPGAs
R. Nepomuceno, R. Sterle, G. Valarini, M. Pereira +2
2021-03-23
Hardware Architecture · Computer Science
Design space exploration for image processing architectures on FPGA targets
Chandrajit Pal, Avik Kotal, Asit Samanta, Amlan Chakrabarti +1
2014-04-16
Distributed, Parallel, and Cluster Computing · Computer Science
Efficient Matrix Factorization on Heterogeneous CPU-GPU Systems
Yuanhang Yu, Dong Wen, Ying Zhang, Xiaoyang Wang +2
2020-06-30
Distributed, Parallel, and Cluster Computing · Computer Science
Flexible Communication Avoiding Matrix Multiplication on FPGA with High-Level Synthesis
Johannes de Fine Licht, Grzegorz Kwasniewski, Torsten Hoefler
2021-01-26