English
Related papers

Related papers: A survey on FPGA-based accelerator for ML models

200 papers

Modern multicore systems are migrating from homogeneous systems to heterogeneous systems with accelerator-based computing in order to overcome the barriers of performance and power walls. In this trend, FPGA-based accelerators are becoming…

Hardware Architecture · Computer Science 2020-09-04 Zhe Lin , Sharad Sinha , Hao Liang , Liang Feng , Wei Zhang

In recent years deep learning algorithms have shown extremely high performance on machine learning tasks such as image classification and speech recognition. In support of such applications, various FPGA accelerator architectures have been…

Machine Learning · Computer Science 2017-05-09 Xinyu Zhang , Srinjoy Das , Ojash Neopane , Ken Kreutz-Delgado

Hardware accelerators are essential for achieving low-latency, energy-efficient inference in edge applications like image recognition. Spiking Neural Networks (SNNs) are particularly promising due to their event-driven and temporally sparse…

Neural and Evolutionary Computing · Computer Science 2026-02-25 Alessio Caviglia , Filippo Marostica , Alessio Carpegna , Alessandro Savino , Stefano Di Carlo

FPGA-based heterogeneous architectures provide programmers with the ability to customize their hardware accelerators for flexible acceleration of many workloads. Nonetheless, such advantages come at the cost of sacrificing programmability.…

Hardware Architecture · Computer Science 2018-07-05 Jason Cong , Zhenman Fang , Yuchen Hao , Peng Wei , Cody Hao Yu , Chen Zhang , Peipei Zhou

Dynamic Graph Neural Networks (DGNNs) are becoming increasingly popular due to their effectiveness in analyzing and predicting the evolution of complex interconnected graph-based systems. However, hardware deployment of DGNNs still remains…

Hardware Architecture · Computer Science 2023-04-17 Hanqiu Chen , Cong Hao

As machine learning (ML) is increasingly implemented in hardware to address real-time challenges in scientific applications, the development of advanced toolchains has significantly reduced the time required to iterate on various designs.…

We describe an efficient FPGA implementation for the exponentiation of large matrices. The research is related to an algorithm for constructing uniformly distributed linear recurring sequences. The design utilizes the special properties of…

Data Structures and Algorithms · Computer Science 2015-03-19 T. Herendi , R. Major

With their widespread availability, FPGA-based accelerators cards have become an alternative to GPUs and CPUs to accelerate computing in applications with certain requirements (like energy efficiency) or properties (like fixed-point…

Hardware Architecture · Computer Science 2022-10-20 Tom Vander Aa , Tom Haber , Thomas J. Ashby , Roel Wuyts , Wilfried Verachtert

Artificial intelligence (AI) methods have become critical in scientific applications to help accelerate scientific discovery. Large language models (LLMs) are being considered as a promising approach to address some of the challenging…

Custom hardware accelerators for Deep Neural Networks are increasingly popular: in fact, the flexibility and performance offered by FPGAs are well-suited to the computational effort and low latency constraints required by many image…

Hardware Architecture · Computer Science 2021-03-25 Serena Curzel , Nicolò Ghielmetti , Michele Fiorito , Fabrizio Ferrandi

This paper presents the High-Performance computing efforts with FPGA for the accelerated pulsar/transient search for the SKA. Case studies are presented from within SKA and pathfinder telescopes highlighting future opportunities. It reviews…

In this community review report, we discuss applications and techniques for fast machine learning (ML) in science -- the concept of integrating power ML methods into the real-time experimental data processing loop to accelerate scientific…

Machine Learning · Computer Science 2023-02-07 Allison McCarn Deiana , Nhan Tran , Joshua Agar , Michaela Blott , Giuseppe Di Guglielmo , Javier Duarte , Philip Harris , Scott Hauck , Mia Liu , Mark S. Neubauer , Jennifer Ngadiuba , Seda Ogrenci-Memik , Maurizio Pierini , Thea Aarrestad , Steffen Bahr , Jurgen Becker , Anne-Sophie Berthold , Richard J. Bonventre , Tomas E. Muller Bravo , Markus Diefenthaler , Zhen Dong , Nick Fritzsche , Amir Gholami , Ekaterina Govorkova , Kyle J Hazelwood , Christian Herwig , Babar Khan , Sehoon Kim , Thomas Klijnsma , Yaling Liu , Kin Ho Lo , Tri Nguyen , Gianantonio Pezzullo , Seyedramin Rasoulinezhad , Ryan A. Rivera , Kate Scholberg , Justin Selig , Sougata Sen , Dmitri Strukov , William Tang , Savannah Thais , Kai Lukas Unger , Ricardo Vilalta , Belinavon Krosigk , Thomas K. Warburton , Maria Acosta Flechas , Anthony Aportela , Thomas Calvet , Leonardo Cristella , Daniel Diaz , Caterina Doglioni , Maria Domenica Galati , Elham E Khoda , Farah Fahim , Davide Giri , Benjamin Hawks , Duc Hoang , Burt Holzman , Shih-Chieh Hsu , Sergo Jindariani , Iris Johnson , Raghav Kansal , Ryan Kastner , Erik Katsavounidis , Jeffrey Krupa , Pan Li , Sandeep Madireddy , Ethan Marx , Patrick McCormack , Andres Meza , Jovan Mitrevski , Mohammed Attia Mohammed , Farouk Mokhtar , Eric Moreno , Srishti Nagu , Rohin Narayan , Noah Palladino , Zhiqiang Que , Sang Eon Park , Subramanian Ramamoorthy , Dylan Rankin , Simon Rothman , Ashish Sharma , Sioni Summers , Pietro Vischia , Jean-Roch Vlimant , Olivia Weng

While FPGA accelerator boards and their respective high-level design tools are maturing, there is still a lack of multi-FPGA applications, libraries, and not least, benchmarks and reference implementations towards sustained HPC usage of…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-03-01 Marius Meyer , Tobias Kenter , Christian Plessl

We present a compilation flow for the generation of CNN inference accelerators on FPGAs. The flow translates a frozen model into OpenCL kernels with the TVM compiler and uses the Intel OpenCL SDK to compile to an FPGA bitstream. We improve…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-03-09 Seung-Hun Chung , Tarek S. Abdelrahman

Convolutional Neural Networks (CNNs) have gained significant traction in the field of machine learning, particularly due to their high accuracy in visual recognition. Recent works have pushed the performance of GPU implementations of CNNs…

Computer Vision and Pattern Recognition · Computer Science 2016-10-03 Roberto DiCecco , Griffin Lacey , Jasmina Vasiljevic , Paul Chow , Graham Taylor , Shawki Areibi

Implementing convolutional neural networks (CNNs) on field-programmable gate arrays (FPGAs) has emerged as a promising alternative to GPUs, offering lower latency, greater power efficiency and greater flexibility. However, this development…

Hardware Architecture · Computer Science 2025-10-21 Philippe Magalhães , Virginie Fresse , Benoît Suffran , Olivier Alata

Large Language Models (LLMs) have propelled groundbreaking advancements across several domains and are commonly used for text generation applications. However, the computational demands of these complex models pose significant challenges,…

As the complexity of deep learning (DL) models increases, their compute requirements increase accordingly. Deploying a Convolutional Neural Network (CNN) involves two phases: training and inference. With the inference task typically taking…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-06-25 Diederik Adriaan Vink , Aditya Rajagopal , Stylianos I. Venieris , Christos-Savvas Bouganis

Particle Accelerators are high power complex machines. To ensure uninterrupted operation of these machines, thousands of pieces of equipment need to be synchronized, which requires addressing many challenges including design, optimization…

Machine Learning · Computer Science 2025-04-08 Kishansingh Rajput , Sen Lin , Auralee Edelen , Willem Blokland , Malachi Schram

FPGA-based data processing in datacenters is increasing in popularity due to the demands of modern workloads and the ensuing necessity for specialization in hardware. Driven by this trend, vendors are rapidly adapting reconfigurable devices…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-04-06 Kaan Kara , Christoph Hagleitner , Dionysios Diamantopoulos , Dimitris Syrivelis , Gustavo Alonso