Machine Learning · Computer Science
A Length Adaptive Algorithm-Hardware Co-design of Transformer on FPGA Through Sparse Attention and Dynamic Pipelining
Hongwu Peng, Shaoyi Huang, Shiyang Chen, Bingbing Li +7
2022-08-23
Hardware Architecture · Computer Science
CPSAA: Accelerating Sparse Attention using Crossbar-based Processing-In-Memory Architecture
Huize Li, Hai Jin, Long Zheng, Yu Huang +6
2023-10-10
Hardware Architecture · Computer Science
SpiDR: A Reconfigurable Digital Compute-in-Memory Spiking Neural Network Accelerator for Event-based Perception
Deepika Sharma, Shubham Negi, Trishit Dutta, Amogh Agrawal +1
2024-11-06
Computer Vision and Pattern Recognition · Computer Science
Generalized Neighborhood Attention: Multi-dimensional Sparse Attention at the Speed of Light
Ali Hassani, Fengzhe Zhou, Aditya Kane, Jiannan Huang +12
2025-04-24
Computation and Language · Computer Science
Native Sparse Attention: Hardware-Aligned and Natively Trainable Sparse Attention
Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo +11
2025-02-28
Hardware Architecture · Computer Science
Sparsity-Aware Hardware-Software Co-Design of Spiking Neural Networks: An Overview
Ilkin Aliyev, Kama Svoboda, Tosiron Adegbija, Jean-Marc Fellous
2024-08-27
Systems and Control · Electrical Eng. & Systems
An Efficient Hardware Accelerator for Structured Sparse Convolutional Neural Networks on FPGAs
Chaoyang Zhu, Kejie Huang, Shuyuan Yang, Ziqi Zhu +2
2020-01-08
Distributed, Parallel, and Cluster Computing · Computer Science
Adaptive Elastic Training for Sparse Deep Learning on Heterogeneous Multi-GPU Servers
Yujing Ma, Florin Rusu, Kesheng Wu, Alexander Sim
2021-10-15
Distributed, Parallel, and Cluster Computing · Computer Science
AsyncSparse: Accelerating Sparse Matrix-Matrix Multiplication on Asynchronous GPU Architectures
Jie Liu, Huanzhi Pu, Zhiru Zhang
2026-04-21
Hardware Architecture · Computer Science
Enabling Flexibility for Sparse Tensor Acceleration via Heterogeneity
Eric Qin, Raveesh Garg, Abhimanyu Bambhaniya, Michael Pellauer +4
2022-01-25
Hardware Architecture · Computer Science
Sense: Model Hardware Co-design for Accelerating Sparse CNN on Systolic Array
Wenhao Sun, Deng Liu, Zhiwei Zou, Wendi Sun +2
2022-09-26
Distributed, Parallel, and Cluster Computing · Computer Science
Batched Sparse Matrix Multiplication for Accelerating Graph Convolutional Networks
Yusuke Nagasaka, Akira Nukada, Ryosuke Kojima, Satoshi Matsuoka
2019-03-28
Machine Learning · Computer Science
Sparse Attention Acceleration with Synergistic In-Memory Pruning and On-Chip Recomputation
Amir Yazdanbakhsh, Ashkan Moradifirouzabadi, Zheng Li, Mingu Kang
2022-09-02
Distributed, Parallel, and Cluster Computing · Computer Science
Design Principles for Sparse Matrix Multiplication on the GPU
Carl Yang, Aydin Buluc, John D. Owens
2018-06-13
Distributed, Parallel, and Cluster Computing · Computer Science
RSH-SpMM: A Row-Structured Hybrid Kernel for Sparse Matrix-Matrix Multiplication on GPUs
Aiying Li, Jingwei Sun, Han Li, Wence Ji +1
2026-03-11
Programming Languages · Computer Science
SparseAuto: An Auto-Scheduler for Sparse Tensor Computations Using Recursive Loop Nest Restructuring
Adhitha Dias, Logan Anderson, Kirshanthan Sundararajah, Artem Pelenitsyn +1
2024-08-20
Distributed, Parallel, and Cluster Computing · Computer Science
FSA: An Alternative Efficient Implementation of Native Sparse Attention Kernel
Ran Yan, Youhe Jiang, Zhuoming Chen, Haohui Mai +2
2025-10-14