English
Related papers

Related papers: Development of Lattice QCD Tool Kit on Cell Broadb…

200 papers

Existing numerical optimizers deployed in quantum compilers use expensive $\mathcal{O}(4^n)$ matrix-matrix operations. Inspired by recent advances in quantum machine learning (QML), QFactor-Sample replaces matrix-matrix operations with…

Quantum Physics · Physics 2024-08-21 Alon Kukliansky , Lukasz Cincio , Ed Younis , Costin Iancu

In the Noisy Intermediate Scale Quantum (NISQ) era, finding implementations of quantum algorithms that minimize the number of expensive and error prone multi-qubit gates is vital to ensure computations produce meaningful outputs. Unitary…

Quantum Physics · Physics 2023-06-12 Mathias Weiden , Ed Younis , Justin Kalloor , John Kubiatowicz , Costin Iancu

Quantum devices in the Noisy Intermediate-Scale Quantum (NISQ) era are limited by high error rates and short decoherence times. Typically, compiler optimisations have provided solutions at the gate level. Alternatively, we exploit the…

Quantum Physics · Physics 2023-11-16 Lilian Hunt Alan Robertson

We present a sub-matrix update algorithm for the continuous-time auxiliary field method that allows the simulation of large lattice and impurity problems. The algorithm takes optimal advantage of modern CPU architectures by consistently…

Strongly Correlated Electrons · Physics 2011-05-09 Emanuel Gull , Peter Staar , Sebastian Fuchs , Phani Nukala , Michael S. Summers , Thomas Pruschke , Thomas Schulthess , Thomas Maier

The yield of physical qubits fabricated in the laboratory is much lower than that of classical transistors in production semiconductor fabrication. Actual implementations of quantum computers will be susceptible to loss in the form of…

Quantum Physics · Physics 2018-01-24 Shota Nagayama , Austin G. Fowler , Dominic Horsman , Simon J. Devitt , Rodney Van Meter

A previously introduced multi-boson technique for the simulation of QCD with dynamical quarks is described and some results of first test runs on a $6^3\times12$ lattice with Wilson quarks and gauge group SU(2) are reported.

High Energy Physics - Lattice · Physics 2009-10-22 B. Bunk , K. Jansen , B. Jegerlehner , M. Lüscher , H. Simma , R. Sommer

The acceleration of deep-learning kernels in hardware relies on matrix multiplications that are executed efficiently on Systolic Arrays (SA). To effectively trade off deep-learning training/inference quality with hardware cost, SA…

Hardware Architecture · Computer Science 2023-09-11 D. Filippas , C. Peltekis , G. Dimitrakopoulos , C. Nicopoulos

Current PC processors are equipped with vector processing units and have other advanced features that can be used to accelerate lattice QCD programs. Clusters of PCs with a high-bandwidth network thus become powerful and cost-effective…

High Energy Physics - Lattice · Physics 2007-05-23 Martin Lüscher

We report on coding and performance of our polynomial hybrid Monte Carlo program on the Earth Simulator. At present the entire program achieves 25--40% efficiency. An analysis of overheads shows that a tuning of inter-node communications is…

High Energy Physics - Lattice · Physics 2009-11-10 S. Aoki , K. -I. Ishikawa , Y. Iwasaki , K. Kanaya , T. Kaneko , Y. Kuramashi , N. Tsutsui , A. Ukawa , T. Yoshie

In recent years, the fervent demand for computational power across various domains has prompted hardware manufacturers to introduce specialized computing hardware aimed at enhancing computational capabilities. Particularly, the utilization…

Numerical Analysis · Mathematics 2024-03-12 Hongyaoxing Gu

We use lattice QCD to calculate the B-mixing hadronic matrix elements for a basis of effective four-quark operators that spans the space of all possible contributions in, and beyond, the Standard Model. We present results for the…

High Energy Physics - Phenomenology · Physics 2015-06-11 C. M. Bouchard

Matrix multiplication is the bedrock in Deep Learning inference application. When it comes to hardware acceleration on edge computing devices, matrix multiplication often takes up a great majority of the time. To achieve better performance…

Machine Learning · Computer Science 2021-10-12 Yuyang Zhang , Dik Hin Leung , Min Guo , Yijia Xiao , Haoyue Liu , Yunfei Li , Jiyuan Zhang , Guan Wang , Zhen Chen

In this paper we develop the first fine-grained rounding error analysis of finite element (FE) cell kernels and assembly. The theory includes mixed-precision implementations and accounts for hardware-acceleration via matrix multiplication…

Numerical Analysis · Mathematics 2024-10-17 M. Croci , G. N. Wells

The devices designed for the Internet-of-Things encompass a large variety of distinct processor architectures, forming a highly heterogeneous zoo. In order to tackle this, we employ a simulator to estimate the performance of the…

Hardware Architecture · Computer Science 2024-03-13 Cristian Ramírez , Adrián Castelló , Héctor Martínez , Enrique S. Quintana-Ortí

Extensions to the C++ implementation of the QCD Data Parallel Interface are provided enabling acceleration of expression evaluation on NVIDIA GPUs. Single expressions are off-loaded to the device memory and execution domain leveraging the…

High Energy Physics - Lattice · Physics 2011-11-24 Frank Winter

The performance of any elliptic curve cryptography hardware accelerator significantly relies on the efficiency of the underlying point multiplication (PM) architecture. This article presents a hardware implementation of field-programmable…

Existing binary Transformers are promising in edge deployment due to their compact model size, low computational complexity, and considerable inference accuracy. However, deploying binary Transformers faces challenges on prior processors…

Hardware Architecture · Computer Science 2024-07-16 Yuhao Ji , Chao Fang , Zhongfeng Wang

The rise of exascale supercomputers has fueled competition among GPU vendors, driving lattice QCD developers to write code that supports multiple APIs. Moreover, new developments in algorithms and physics research require frequent updates…

We present the first lattice calculation of the matrix element of the electromagnetic operator <pi0|Q+|K0>, where Q+ = (Q_d e/16 pi^2)* (\bar s_L sigma{mu,nu} F{mu,nu} d_R + \bar s_R sigma{mu,nu} F{mu,nu} d_L). This matrix element plays an…

High Energy Physics - Phenomenology · Physics 2015-06-25 D. Becirevic , V. Lubicz , G. Martinelli , F. Mescia

As users and developers, we are witnessing the opening of a new computing scenario: the introduction of hybrid processors into a single die, such as an accelerated processing unit (APU) processor, and the plug-and-play of additional…

Mathematical Software · Computer Science 2012-05-15 Paolo D'Alberto