Related papers: Accelerating Twisted Mass LQCD with QPhiX
We compare the behavior of overlap fermions, which are chirally invariant, and of Wilson twisted mass fermions at full twist in the approach to the chiral limit. Our quenched simulations reveal that with both formulations of lattice…
For years, SIMD/vector units have enhanced the capabilities of modern CPUs in High-Performance Computing (HPC) and mobile technology. Typical commercially-available SIMD units process up to 8 double-precision elements with one instruction.…
Accelerating Human Action Recognition (HAR) efficiently for real-time surveillance and robotic systems on edge chips remains a challenging research field, given its high computational and memory requirements. This paper proposed an…
This work presents a lattice quantum chromodynamics (QCD) calculation of the nonperturbative Collins-Soper kernel, which describes the rapidity evolution of quark transverse-momentum-dependent parton distribution functions. The kernel is…
In this paper, we present the details of our multi-node GPU-FFT library, as well its scaling on Selene HPC system. Our library employs slab decomposition for data division and MPI for communication among GPUs. We performed GPU-FFT on…
Data movement overheads increase the inference latency of state-of-the-art large language models (LLMs). These models commonly use the bfloat16 (BF16) format for stable training. Floating-point standards allocate eight bits to the exponent,…
Time-domain nonvolatile in-memory computing (TD-nvIMC) offers a promising pathway to reduce data movement and improve energy efficiency by encoding computation in delay rather than voltage or current. This work presents a fully integrated…
Fabrication of quantum processors in advanced 300 mm wafer-scale complementary metal-oxide-semiconductor (CMOS) foundries provides a unique scaling pathway towards commercially viable quantum computing with potentially millions of qubits on…
Chip-scale, high-energy optical pulse generation is becoming increasingly important as we expand activities into hard to reach areas such as space and deep ocean. Q-switching of the laser cavity is the best known technique for generating…
Important memory-bound kernels, such as linear algebra, convolutions, and stencils, rely on SIMD instructions as well as optimizations targeting improved vectorized data traversal and data re-use to attain satisfactory performance. On on…
This paper is a slightly modified and reduced version of the proposal of the {\bf apeNEXT} project, which was submitted to DESY and INFN in spring 2000. .It presents the basic motivations and ideas of a next generation lattice QCD (LQCD)…
In this paper, we propose a high-performance RISC-V soft processor with an efficient fetch unit supporting the compressed instructions targeting on FPGA. The compressed instruction extension in RISC-V can reduce the program size by about…
We have studied the light hadron spectrum and decay constants for quenched QCD at beta=6.2 on a 24^3x48 lattice. We compare the results obtained using a nearest-neighbour O(a)-improved ("clover") fermion action with those obtained using the…
Whilst numerous areas of computing have adopted the RISC-V Instruction Set Architecture (ISA) wholesale in recent years, it is yet to become widespread in HPC. RISC-V accelerators offer a compelling option where the HPC community can…
BERT is the most recent Transformer-based model that achieves state-of-the-art performance in various NLP tasks. In this paper, we investigate the hardware acceleration of BERT on FPGA for edge computing. To tackle the issue of huge…
The growing concerns regarding energy consumption and privacy have prompted the development of AI solutions deployable on the edge, circumventing the substantial CO2 emissions associated with cloud servers and mitigating risks related to…
The emerging trend of deploying complex algorithms, such as Deep Neural Networks (DNNs), increasingly poses strict memory and energy efficiency requirements on Internet-of-Things (IoT) end-nodes. Mixed-precision quantization has been…
Hardware accelerations of deep learning systems have been extensively investigated in industry and academia. The aim of this paper is to achieve ultra-high energy efficiency and performance for hardware implementations of deep neural…
We report the performance of Wilson and Domainwall Kernels on a new Intel Xeon Phi Knights Landing based machine named Oakforest-PACS, which is co-hosted by University of Tokyo and Tsukuba University and is currently fastest in Japan. This…
We present preliminary results for the spectrum and decay matrix elements for heavy-light and heavy-heavy mesons, obtained on the 64-node Meiko Computing Surface at the University of Edinburgh. Quark propagators are computed with an…