English
Related papers

Related papers: Optimization of Lattice QCD codes for the AMD Opte…

200 papers

In this paper we propose a new design criterion and a new class of unitary signal constellations for differential space-time modulation for multiple-antenna systems over Rayleigh flat-fading channels with unknown fading coefficients.…

Information Theory · Computer Science 2008-05-12 Xinjia Chen , Kemin Zhou , Jorge Aravena

In this paper, we propose a novel reordering scheme to improve the performance of a Laplacian Mesh Smoothing (LMS). While the Laplacian smoothing algorithm is well optimized and studied, we show how a simple reordering of the vertices of…

Numerical Analysis · Computer Science 2016-06-03 Guillaume Aupy , JeongHyung Park , Padma Raghavan

The architecture and capabilities of the computers currently in use for large-scale lattice QCD calculations are described and compared. Based on this present experience, possible future directions are discussed.

High Energy Physics - Lattice · Physics 2015-06-25 Norman H. Christ

This paper presents a new class of sparse superposition codes for low-rates and short-packet communications over the additive white Gaussian noise channel. The new code is orthogonal sparse superposition (OSS) code. A codeword of OSS codes…

Information Theory · Computer Science 2020-11-24 Yunseo Nam , Jeonghun Park , Songnam Hong , Namyoon Lee

Solving discretized versions of the Dirac equation represents a large share of execution time in lattice Quantum Chromodynamics (QCD) simulations. Many high-performance computing (HPC) clusters use graphics processing units (GPUs) to offer…

High Energy Physics - Lattice · Physics 2024-07-02 Tilmann Matthaei

Quantum error correction (QEC) is essential for quantum computing to mitigate the effect of errors on qubits, and surface code (SC) is one of the most promising QEC methods. Decoding SCs is the most computational expensive task in the…

Quantum Physics · Physics 2022-09-02 Yosuke Ueno , Masaaki Kondo , Masamitsu Tanaka , Yasunari Suzuki , Yutaka Tabuchi

Quantum error correction (QEC) for fault-tolerant quantum computing requires a balanced decoding solution that offers high performance, low complexity, and low latency. However, the de facto standard, belief propagation (BP) combined with…

Quantum Physics · Physics 2026-05-04 Hee-Youl Kwak , Seong-Joon Park , Hyunwoo Jung , Jeongseok Ha , Jae-Won Kim

In this work, we present an optimal mapper for OFDM with index modulation (OFDM-IM). By optimal we mean the mapper achieves the lowest possible asymptotic computational complexity (CC) when the spectral efficiency (SE) gain over OFDM…

Signal Processing · Electrical Eng. & Systems 2020-04-07 Saulo Queiroz , João P. Vilela , Edmundo Monteiro

We publish an extension of openQCD-1.6 with AVX-512 vector instructions using Intel intrinsics. Recent Intel processors support extended instruction sets with operations on 512-bit wide vectors, increasing both the capacity for floating…

High Energy Physics - Lattice · Physics 2018-11-22 Ed Bennett , Mark Dawson , Michele Mesiti , Jarno Rantaharju

FDTD codes, such as Sophie developed at CEA/DAM, no longer take advantage of the processor's increased computing power, especially recently with the raising multicore technology. This is rooted in the fact that low order numerical schemes…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-01-22 Olivier Cessenat

This paper presents a novel, non-standard set of vector instruction types for exploring custom SIMD instructions in a softcore. The new types allow simultaneous access to a relatively high number of operands, reducing the instruction count…

Hardware Architecture · Computer Science 2021-06-15 Philippos Papaphilippou , Paul H. J. Kelly , Wayne Luk

SSE (streaming SIMD extensions) and AVX (advanced vector extensions) are SIMD (single instruction multiple data streams) instruction sets supported by recent CPUs manufactured in Intel and AMD. This SIMD programming allows parallel…

High Energy Physics - Lattice · Physics 2013-11-05 Hwancheol Jeong , Sunghoon Kim , Weonjong Lee , Seok-Ho Myung

As the first kind of forward error correction (FEC) codes that achieve channel capacity, polar codes have attracted much research interest recently. Compared with other popular FEC codes, polar codes decoded by list successive cancellation…

Signal Processing · Electrical Eng. & Systems 2018-07-02 ChenYang Xia , Ji Chen , YouZhe Fan , Chi-ying Tsui , Jie Jin , Hui Shen , Bin Li

Sparse superposition codes are a recent class of codes introduced by Barron and Joseph for efficient communication over the AWGN channel. With an appropriate power allocation, these codes have been shown to be asymptotically…

Information Theory · Computer Science 2018-03-19 Adam Greig , Ramji Venkataramanan

The speed, bandwidth and cost characteristics of today's PC graphics cards make them an attractive target as general purpose computational platforms. High performance can be achieved also for lattice simulations but the actual…

High Energy Physics - Lattice · Physics 2008-11-26 Gyozo I. Egri , Zoltan Fodor , Christian Hoelbling , Sandor D. Katz , Daniel Nogradi , Kalman K. Szabo

Long-latency load requests continue to limit the performance of high-performance processors. To increase the latency tolerance of a processor, architects have primarily relied on two key techniques: sophisticated data prefetchers and large…

Hardware Architecture · Computer Science 2022-10-03 Rahul Bera , Konstantinos Kanellopoulos , Shankar Balachandran , David Novo , Ataberk Olgun , Mohammad Sadrosadati , Onur Mutlu

Creating high performance implementations of deep learning primitives on CPUs is a challenging task. Multiple considerations including multi-level cache hierarchy, and wide SIMD units of CPU platforms influence the choice of program…

Programming Languages · Computer Science 2021-04-13 Sanket Tavarageri , Gagandeep Goyal , Sasikanth Avancha , Bharat Kaul , Ramakrishna Upadrasta

We report an implementation of a code for SU(3) matrix multiplication on Cell/B.E., which is a part of our project, Lattice Tool Kit on Cell/B.E.. On QS20, the speed of the matrix multiplication on SPE in single precision is 227GFLOPS and…

High Energy Physics - Lattice · Physics 2012-03-16 Shinji Motok , i Yoshiyuki Nakagawa , Keitaro Nagata , Koichi Hashimoto , Kiyoshi Mizumaru , Atsushi Nakamura

Implementations of measurement kernels in high-level Lattice QCD frameworks enable rapid prototyping, but can leave hardware capabilities significantly underutilized. This is an acceptable tradeoff if the time spent in unoptimized routines…

High Energy Physics - Lattice · Physics 2022-11-30 Phuong Nguyen , Ben Hörz

We study the implementation of the even-odd Wilson fermion matrix for lattice QCD simulations on the A64FX architecture. Efficient coding of the stencil operation is investigated for two-dimensional packing to SIMD vectors. We measure the…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-03-16 Issaku Kanamori , Keigo Nitadori , Hideo Matsufuru
‹ Prev 1 3 4 5 6 7 10 Next ›