English
Related papers

Related papers: apenext: A Multi-Tflops LQCD Computing Project

200 papers

In October, 2016, the US Department of Energy launched the Exascale Computing Project, which aims to deploy exascale computing resources for science and engineering in the early 2020's. The project brings together application teams,…

High Energy Physics - Lattice · Physics 2018-04-18 Richard Brower , Norman Christ , Carleton DeTar , Robert Edwards , Paul Mackenzie

Recent progress in Lattice QCD is highlighted. After a brief introduction to the methodology of lattice computations the presentation focuses on three main topics: Hadron Spectroscopy, Hadron Structure and Lattice Flavor Physics. In each…

High Energy Physics - Lattice · Physics 2013-01-10 Stephan Durr

In this paper we introduce Epiphany as a high-performance energy-efficient manycore architecture suitable for real-time embedded systems. This scalable architecture supports floating point operations in hardware and achieves 50 GFLOPS/W in…

Hardware Architecture · Computer Science 2014-12-18 Andreas Olofsson , Tomas Nordström , Zain Ul-Abdin

As large language models (LLMs) become increasingly powerful, the sequential nature of autoregressive generation creates a fundamental throughput bottleneck that limits the practical deployment. While Multi-Token Prediction (MTP) has…

Machine Learning · Computer Science 2025-09-24 Yuxuan Cai , Xiaozhuan Liang , Xinghua Wang , Jin Ma , Haijin Liang , Jinwen Luo , Xinyu Zuo , Lisheng Duan , Yuyang Yin , Xi Chen

Hardware acceleration for dilated and transposed convolution enables real time execution of related tasks like segmentation, but current designs are specific for these convolutional types or suffer from complex control for reconfigurable…

Hardware Architecture · Computer Science 2022-05-05 Kuo-Wei Chang , Tian-Sheuan Chang

While the HPC community is working towards the development of the first Exaflop computer (expected around 2020), after reaching the Petaflop milestone in 2008 still only few HPC applications are able to fully exploit the capabilities of…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-06-10 Erika Abraham , Costas Bekas , Ivona Brandic , Samir Genaim , Einar Broch Johnsen , Ivan Kondov , Sabri Pllana , Achim Streit

We propose without loss of generality strategies to achieve a high-throughput FPGA-based architecture for a QC-LDPC code based on a circulant-1 identity matrix construction. We present a novel representation of the parity-check matrix (PCM)…

Hardware Architecture · Computer Science 2015-05-12 Swapnil Mhaske , Hojin Kee , Tai Ly , Ahsan Aziz , Predrag Spasojevic

In this paper, we propose LoopLynx, a scalable dataflow architecture for efficient LLM inference that optimizes FPGA usage through a hybrid spatial-temporal design. The design of LoopLynx incorporates a hybrid temporal-spatial architecture,…

Hardware Architecture · Computer Science 2025-04-15 Jianing Zheng , Gang Chen

Reconciliation is a crucial procedure in post-processing of Quantum Key Distribution (QKD), which is used for correcting the error bits in sifted key strings. Although most studies about reconciliation of QKD focus on how to improve the…

Quantum Physics · Physics 2019-03-26 Haokun Mao , Qiong Li , Qi Han , Hong Guo

High-performance computing systems are more and more often based on accelerators. Computing applications targeting those systems often follow a host-driven approach in which hosts offload almost all compute-intensive sections of the code…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-05-15 E. Calore , A. Gabbana , S. F. Schifano , R. Tripiccione

Optimizing deep learning models is generally performed in two steps: (i) high-level graph optimizations such as kernel fusion and (ii) low level kernel optimizations such as those found in vendor libraries. This approach often leaves…

Machine Learning · Computer Science 2021-03-08 Pratik Fegade , Tianqi Chen , Phillip B. Gibbons , Todd C. Mowry

L-band Digital Aeronautical Communication System (LDACS) aims to exploit vacant spectrum in L-band via spectrum sharing, and orthogonal frequency division multiplexing (OFDM) is the currently accepted LDACS waveform. Recently, various works…

Signal Processing · Electrical Eng. & Systems 2021-01-19 N. Agrawal , A. Ambede , S. J. Darak , A. P. Vinod , A. S. Madhukumar

The search for new physics requires a joint experimental and theoretical effort. Lattice QCD is already an essential tool for obtaining precise model-free theoretical predictions of the hadronic processes underlying many key experimental…

This is the report of the Computing Frontier working group on Lattice Field Theory prepared for the proceedings of the 2013 Community Summer Study ("Snowmass"). We present the future computing needs and plans of the U.S. lattice gauge…

High Energy Physics - Lattice · Physics 2013-10-25 T. Blum , R. S. Van de Water , D. Holmgren , R. Brower , S. Catterall , N. Christ , A. Kronfeld , J. Kuti , P. Mackenzie , E. T. Neil , S. R. Sharpe , R. Sugar

Fast Fourier Transform (FFT) is an essential tool in scientific and engineering computation. The increasing demand for mixed-precision FFT has made it possible to utilize half-precision floating-point (FP16) arithmetic for faster speed and…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-04-26 Binrui Li , Shenggan Cheng , James Lin

I review the most recent evolutions of the QCD codes on new architectures, with a focus on the performances obtained by the different coding strategies as presented during the Lattice2017 conference.

High Energy Physics - Lattice · Physics 2018-04-18 Antonio Rago

Running from October 2011 to June 2015, the aim of the European project Mont-Blanc has been to develop an approach to Exascale computing based on embedded power-efficient technology. The main goals of the project were to i) build an HPC…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-08-21 Momme Allalen , David Brayford , Daniele Tafani , Volker Weinberg , Bernd Mohr , Dirk Brömmel , Rene Halver , Jan Meinke , Sandipan Mohanty

Mini-batch inference of Graph Neural Networks (GNNs) is a key problem in many real-world applications. Recently, a GNN design principle of model depth-receptive field decoupling has been proposed to address the well-known issue of…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-01-05 Bingyi Zhang , Hanqing Zeng , Viktor Prasanna

Many edge devices employ Recurrent Neural Networks (RNN) to enhance their product intelligence. However, the increasing computation complexity poses challenges for performance, energy efficiency and product development time. In this paper,…

Neural and Evolutionary Computing · Computer Science 2020-10-27 Chao-Yang Kao , Huang-Chih Kuo , Jian-Wen Chen , Chiung-Liang Lin , Pin-Han Chen , Youn-Long Lin

The rapid growth of deep learning has driven exponential increases in model parameters and computational demands. NVIDIA GPUs and their CUDA-based software ecosystem provide robust support for parallel computing, significantly alleviating…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-07-08 Jiaqi Lv , Xufeng He , Yanchen Liu , Xu Dai , Aocheng Shen , Yinghao Li , Jiachen Hao , Jianrong Ding , Yang Hu , Shouyi Yin