Related papers: EPAC: The Last Dance
CLIC is a proposed linear $e^{+}e^{-}$ collider with center-of-mass energies of up to $3\,\textrm{TeV}$. Its main objectives are precise top quark and Higgs boson measurements, as well as searches for Beyond Standard Model physics. To meet…
Real-time, energy-efficient inference on edge devices is essential for graph classification across a range of applications. Hyperdimensional Computing (HDC) is a brain-inspired computing paradigm that encodes input features into…
We propose an optimization approach for determining both hardware and software parameters for the efficient implementation of a (family of) applications called dense stencil computations on programmable GPGPUs. We first introduce a simple,…
We have developed a position-sensitive Parallel Plate Avalanche Counter (PPAC), which has been used as a focal plane detector in the BigRIPS fragment separator and the subsequent RI-beam delivery lines at the RIKEN Nishina Center RI Beam…
We present a parallel implementation of a direct solver for the Poisson's equation on extreme-scale supercomputers with accelerators. We introduce a chunked-pencil decomposition as the domain-decomposition strategy to distribute work among…
The Advanced Encryption Standard (AES) is a widely adopted cryptographic algorithm essential for securing embedded systems and IoT platforms. However, existing AES hardware accelerators often face limitations in performance, energy…
With the ultimate goal of developing a pixel-based readout for a TPC at the ILC, a GridPix readout system consisting of one Timepix3 chip with an integrated amplification grid was embedded in a prototype detector. The performance was…
The Circular Electron Positron Collider (CEPC) is a large international scientific project initiated and hosted by China. It is located in a 100-km circumference underground tunnel. The accelerator complex consists of a linear accelerator…
An Application-Specific Instruction Set Processor(ASIP) is a specialized microprocessor that provides a trade-off between the programmability of a General Purpose Processor (GPP) and the performance and energy-efficiency of dedicated…
An array of Parallel Plate Avalanche Counters (PPAC) for the detection of heavy ions has been developed. The new device, NIFF (Nuclear Instrument for Fission Fragments), consists of four individual detectors and covers $60\%$ of 2$\pi$. It…
Generative model based image lossless compression algorithms have seen a great success in improving compression ratio. However, the throughput for most of them is less than 1 MB/s even with the most advanced AI accelerated chips, preventing…
General Matrix Multiplication (GEMM) is a critical operation underpinning a wide range of applications in high-performance computing (HPC) and artificial intelligence (AI). The emergence of hardware optimized for low-precision arithmetic…
The ePIC collaboration is developing a multidetector system to explore the fundamental properties of the strong interaction at the future Electron-Ion Collider (EIC), to be built at Brookhaven National Laboratory. A key component of the…
This paper presents a methodology for simultaneous heterogeneous computing, named ENEAC, where a quad core ARM Cortex-A53 CPU works in tandem with a preprogrammed on-board FPGA accelerator. A heterogeneous scheduler distributes the tasks…
Whilst RISC-V has grown phenomenally quickly in embedded computing, it is yet to gain significant traction in High Performance Computing (HPC). However, as we move further into the exascale era, the flexibility offered by RISC-V has the…
The slowdown of Moore's law and the power wall necessitates a shift towards finely tunable precision (a.k.a. transprecision) computing to reduce energy footprint. Hence, we need circuits capable of performing floating-point operations on a…
This paper presents SynapticCore-X, a modular and resource-efficient neural processing architecture optimized for deployment on low-cost FPGA platforms. The design integrates a lightweight RV32IMC RISC-V control core with a configurable…
Artificial intelligence necessitates adaptable hardware accelerators for efficient high-throughput million operations. We present pipelined architecture with CORDIC block for linear MAC computations and nonlinear iterative Activation…
We study parallel particle-in-cell (PIC) methods for low-temperature plasmas (LTPs), which discretize kinetic formulations that capture the time evolution of the probability density function of particles as a function of position and…
Data logging at an upgraded KEKB accelerator or the J-PARC facility, currently under commission, requires a high density data acquisition platform with integrated data reduction CPUs. To follow market trends, we have developed a DAQ…