Distributed, Parallel, and Cluster Computing · Computer Science
Evaluating Spatial Accelerator Architectures with Tiled Matrix-Matrix Multiplication
Gordon E. Moon, Hyoukjun Kwon, Geonhwa Jeong, Prasanth Chatarasi +2
2021-06-22
Hardware Architecture · Computer Science
A Flexible Instruction Set Architecture for Efficient GEMMs
Alexandre de Limas Santana, Adrià Armejach, Francesc Martinez, Erich Focht +1
2025-07-08
Machine Learning · Computer Science
A High-Level Compiler Integration Approach for Deep Learning Accelerators Supporting Abstraction and Optimization
Samira Ahmadifarsani, Daniel Mueller-Gritschneder, Ulf Schlichtmann
2025-07-08
Distributed, Parallel, and Cluster Computing · Computer Science
PALM: A Efficient Performance Simulator for Tiled Accelerators with Large-scale Model Training
Jiahao Fang, Huizheng Wang, Qize Yang, Dehao Kong +4
2024-06-07
Distributed, Parallel, and Cluster Computing · Computer Science
Toward matrix multiplication for deep learning inference on the Xilinx Versal
Jie Lei, José Flich, Enrique S. Quintana-Ortí
2023-02-16
Distributed, Parallel, and Cluster Computing · Computer Science
Leveraging Hardware-Aware Computation in Mixed-Precision Matrix Multiply: A Tile-Centric Approach
Qiao Zhang, Rabab Alomairy, Dali Wang, Zhuowei Gu +1
2025-08-21
Hardware Architecture · Computer Science
GEN-Graph: Heterogeneous PIM Accelerator for General Computational Patterns in Graph-based Dynamic Programming
Yanru Chen, Runyang Tian, Zheyu Li, Mahbod Afarin +2
2026-04-20
Distributed, Parallel, and Cluster Computing · Computer Science
Understanding GEMM Performance and Energy on NVIDIA Ada Lovelace: A Machine Learning-Based Analytical Approach
Xiaoteng, Liu, Pavly Halim
2024-11-27
Hardware Architecture · Computer Science
MACO: Exploring GEMM Acceleration on a Loosely-Coupled Multi-core Processor
Bingcai Sui, Junzhong Shen, Caixia Sun, Junhui Wang +2
2024-05-01
Hardware Architecture · Computer Science
A Highly Configurable Hardware/Software Stack for DNN Inference Acceleration
Suvadeep Banerjee, Steve Burns, Pasquale Cocchini, Abhijit Davare +5
2021-12-01
Hardware Architecture · Computer Science
VEGETA: Vertically-Integrated Extensions for Sparse/Dense GEMM Tile Acceleration on CPUs
Geonhwa Jeong, Sana Damani, Abhimanyu Rajeshkumar Bambhaniya, Eric Qin +4
2023-02-24
Distributed, Parallel, and Cluster Computing · Computer Science
A Dense Tensor Accelerator with Data Exchange Mesh for DNN and Vision Workloads
Yu-Sheng Lin, Wei-Chao Chen. Chia-Lin Yang, Shao-Yi Chien
2021-11-29
Data Structures and Algorithms · Computer Science
Stream-K: Work-centric Parallel Decomposition for Dense Matrix-Matrix Multiplication on the GPU
Muhammad Osama, Duane Merrill, Cris Cecka, Michael Garland +1
2023-01-11
Hardware Architecture · Computer Science
MixDiT: Accelerating Image Diffusion Transformer Inference with Mixed-Precision MX Quantization
Daeun Kim, Jinwoo Hwang, Changhun Oh, Jongse Park
2025-04-14
Mathematical Software · Computer Science
Automating the Last-Mile for High Performance Dense Linear Algebra
Richard Michael Veras, Tze Meng Low, Tyler Michael Smith, Robert van de Geijn +1
2017-05-01
Distributed, Parallel, and Cluster Computing · Computer Science
Anatomy of High-Performance GEMM with Online Fault Tolerance on GPUs
Shixun Wu, Yujia Zhai, Jinyang Liu, Jiajun Huang +3
2023-05-03
Hardware Architecture · Computer Science
Exploring the Performance Improvement of Tensor Processing Engines through Transformation in the Bit-weight Dimension of MACs
Qizhe Wu, Huawen Liang, Yuchen Gui, Zhichen Zeng +8
2025-03-11
Hardware Architecture · Computer Science
Performance Analysis of Matrix Multiplication for Deep Learning on the Edge
Cristian Ramírez, Adrián Castelló, Héctor Martínez, Enrique S. Quintana-Ortí
2024-03-13
Hardware Architecture · Computer Science
METRO: A Software-Hardware Co-Design of Interconnections for Spatial DNN Accelerators
Zhao Wang, Jingchen Zhu, Zhe Zhou, Guangyu Sun
2025-02-19
Hardware Architecture · Computer Science
AME-PIM: Can Memory be Your Next Tensor Accelerator?
Emanuele Venieri, Simone Manoni, Alberto Florian, Jaehyun Park +2
2026-05-01
Distributed, Parallel, and Cluster Computing · Computer Science
Accelerating 128-bit Floating-Point Matrix Multiplication on FPGAs
Fumiya Kono, Naohito Nakasato, Maho Nakata
2023-06-12
Hardware Architecture · Computer Science
Efficient Approaches for GEMM Acceleration on Leading AI-Optimized FPGAs
Endri Taka, Dimitrios Gourounas, Andreas Gerstlauer, Diana Marculescu +1
2024-04-18