English
Related papers

Related papers: A computational system for lattice QCD with overla…

200 papers

To unleash the potential of quantum computers, noise effects on qubits' performance must be carefully managed. The decoders responsible for diagnosing noise-induced computational errors must use resources efficiently to enable scaling to…

Modern interconnects often have programmable processors in the network interface that can be utilized to offload communication processing from host CPU. In this paper, we explore different schemes to support collective operations at the…

Distributed, Parallel, and Cluster Computing · Computer Science 2007-05-23 Weikuan Yu , Darius Buntinas , Rich L. Graham , Dhabaleswar K. Panda

The rapid development of programmable network devices and the widespread use of machine learning (ML) in networking have facilitated efficient research into intelligent data plane (IDP). Offloading ML to programmable data plane (PDP)…

Networking and Internet Architecture · Computer Science 2025-06-24 Mai Zhang , Lin Cui , Xiaoquan Zhang , Fung Po Tso , Zhen Zhang , Yuhui Deng , Zhetao Li

CPUs are critical for LLM serving due to their availability, cost efficiency, and edge applicability. However, efficient CPU serving is hindered by conflicting prefill/decode resource demands under non-disaggregated deployment…

Hardware Architecture · Computer Science 2026-04-16 Juntao Zhao , Jiuru Li , Chuan Wu

We calculate the 2-loop partition function of QCD on the lattice, using the Wilson formulation for gluons and the overlap-Dirac operator for fermions. Direct by-products of our result are the 2-loop free energy and average plaquette. Our…

High Energy Physics - Lattice · Physics 2009-11-10 A. Athenodorou , H. Panagopoulos

This paper proposes a composite inner-product computation unit based on left-to-right (LR) arithmetic for the acceleration of convolution neural networks (CNN) on hardware. The efficacy of the proposed L2R-CIPU method has been shown on the…

Hardware Architecture · Computer Science 2024-07-10 Malik Zohaib Nisar , Mohammad Sohail Ibrahim , Muhammad Usman , Jeong-A Lee

We present a new set of QCD codes in both message passing and data parallel versions. The message passing package used is PARMACS, although other packages may be used. Data parallel software is written in High Performance fortran, an…

High Energy Physics - Lattice · Physics 2007-05-23 Nick Stanford

The current status of United States projects pursuing Teraflops-scale computing resources for lattice field theory is discussed. Two projects are in existence at this time: the Multidisciplinary Teraflops Project, incorporating the…

High Energy Physics - Lattice · Physics 2009-10-28 Robert D. Mawhinney

Specialized hardware like application-specific integrated circuits (ASICs) remains the primary accelerator type for cryptographic kernels based on large integer arithmetic. Prior work has shown that commodity and server-class GPUs can…

Cryptography and Security · Computer Science 2025-09-17 Naifeng Zhang , Sophia Fu , Franz Franchetti

Modern graphics hardware is designed for highly parallel numerical tasks and provides significant cost and performance benefits. Graphics hardware vendors are now making available development tools to support general purpose high…

High Energy Physics - Lattice · Physics 2009-01-22 Kipton Barros , Ronald Babich , Richard Brower , Michael A. Clark , Claudio Rebbi

We present results of a quenched QCD simulation with overlap fermions on a lattice of volume V = 16^3X32 at beta=6.0, which corresponds approximatively to a lattice cutoff of 2 GeV and an extension of 1.4 fm. From the two-point correlation…

High Energy Physics - Lattice · Physics 2015-06-25 L. Giusti , C. Hoelbling , C. Rebbi

Convolutional neural networks (CNNs) require high throughput hardware accelerators for real time applications owing to their huge computational cost. Most traditional CNN accelerators rely on single core, linear processing elements (PEs) in…

Hardware Architecture · Computer Science 2020-07-21 Mahmood Azhar Qureshi , Arslan Munir

We study physics at temperatures just above the QCD phase transition (Tc) using chiral (overlap) Fermions in the quenched approximation of lattice QCD. Exact zero modes of the overlap Dirac operator are localized and their frequency of…

High Energy Physics - Lattice · Physics 2011-07-19 Rajiv V. Gavai , Sourendu Gupta , R. Lacaze

Using barebone PC components and NIC's, we construct a linux cluster which has 2-dimensional mesh structure. This cluster has smaller footprint, is less expensive, and use less power compared to conventional linux cluster. Here, we report…

High Energy Physics - Lattice · Physics 2009-11-10 Chang-Yeong Choi , Jeong-Hyun Kim , Seyong Kim

We propose Shuttling-based Distributed Quantum Computing (SDQC), a hybrid architecture that combines the strengths of physical qubit shuttling and distributed quantum computing to enable scalable trapped-ion quantum computing. SDQC performs…

Quantum Physics · Physics 2025-12-03 Seunghyun Baek , Seok-Hyung Lee , Dongmoon Min , Junki Kim

From the overlap lattice quark propagator calculated in the Landau gauge, we determine the quark chiral condensate by fitting operator product expansion formulas to the lattice data. The quark propagators are computed on domain wall fermion…

High Energy Physics - Lattice · Physics 2017-10-25 Chao Wang , Yujiang Bi , Hao Cai , Ying Chen , Ming Gong , Zhaofeng Liu

LINAC 4 is a normal conducting H- linac proposed at CERN to provide a higher proton flux to the CERN accelerator chain. It should replace the existing LINAC 2 as injector to the Proton Synchrotron Booster and can also operate in the future…

Accelerator Physics · Physics 2007-05-23 M. Baylac , J. M. De Conto , E. Froidefond , E. Sargsyan

TCP/IP network stack is irreplaceable for Web services in datacenter front-end servers, and the demand for which is growing rapidly for emerging high concurrency network service applications (including Internet, Internet of Things, mobile…

Networking and Internet Architecture · Computer Science 2022-10-18 WL Zhang , YF Shen , H Song , Zh Zhang , K Liu , Q Huang , MY Chen

Large deep learning models have demonstrated strong ability to solve many tasks across a wide range of applications. Those large models typically require training and inference to be distributed. Tensor parallelism is a common technique…

We have developed a new version of the high-performance J\"ulich universal quantum computer simulator (JUQCS-50) that leverages key features of the GH200 superchips as used in the JUPITER supercomputer, enabling simulations of a 50-qubit…

‹ Prev 1 8 9 10 Next ›