English
Related papers

Related papers: Solving the Dirac equation on QPACE

200 papers

In cryptanalysis, solving the discrete logarithm problem (DLP) is key to assessing the security of many public-key cryptosystems. The index-calculus methods, that attack the DLP in multiplicative subgroups of finite fields, require solving…

Cryptography and Security · Computer Science 2014-12-05 Hamza Jeljeli

We describe the construction of a high performance parallel computer composed of PC components, present some physical results for light hadron and hybrid meson masses from lattice QCD. We also show that the smearing technique is very useful…

High Energy Physics - Lattice · Physics 2007-05-23 Xiang-Qian Luo , Zhong-Hao Mei , Eric B. Gregory , Jie-Chao Yang , Yu-Li Wang , Yin Lin

In lattice QCD the calculation of disconnected quark loops from the trace of the inverse quark matrix has large noise variance. A multilevel Monte Carlo method is proposed for this problem that uses different degree polynomials on a…

High Energy Physics - Lattice · Physics 2024-02-02 Paul Lashomb , Ronald B. Morgan , Travis Whyte , Walter Wilcox

The limited number of qubits per chip remains a critical bottleneck in quantum computing, motivating the use of distributed architectures that interconnect multiple quantum processing units (QPUs). However, executing quantum algorithms…

Quantum Physics · Physics 2026-01-21 Brayden Goldstein-Gelb , Kun Liu , John M. Martyn , Hengyun , Zhou , Yongshan Ding , Yuan Liu

Leveraging Trace Theory, we investigate the efficient parallelization of direct solvers for large linear equation systems. Our focus lies on a multi-frontal algorithm, and we present a methodology for achieving near-optimal scheduling on…

Numerical Analysis · Mathematics 2023-06-16 Jan Trynda , Maciej Woźniak , Sergio Rojas

OpenMP parallelization of multiple precision Taylor series method is proposed. A very good parallel performance scalability and parallel efficiency inside one computation node of a CPU-cluster is observed. We explain the details of the…

Mathematical Software · Computer Science 2019-08-27 S. Dimova , I. Hristov , R. Hristova , I. Puzynin , T. Puzynina , Z. Sharipov , N. Shegunov , Z. Tukhliev

Large language models (LLMs) suffer from low efficiency as the mismatch between the requirement of auto-regressive decoding and the design of most contemporary GPUs. Specifically, billions to trillions of parameters must be loaded to the…

Computation and Language · Computer Science 2024-05-02 Bin Xiao , Chunan Shi , Xiaonan Nie , Fan Yang , Xiangwei Deng , Lei Su , Weipeng Chen , Bin Cui

A hybrid MPI+OpenMP strategy for parallelizing multiple precision Taylor series method is proposed, realized and tested. To parallelize the algorithm we combine MPI and OpenMP parallel technologies together with GMP library (GNU miltiple…

Mathematical Software · Computer Science 2023-11-27 I. Hristov , R. Hristova , S. Dimova , P. Armyanov , N. Shegunov , I. Puzynin , T. Puzynina , Z. Sharipov , Z. Tukhliev

Bloom filters are a fundamental data structure for approximate membership queries, with applications ranging from data analytics to databases and genomics. Several variants have been proposed to accommodate parallel architectures. GPUs,…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-18 Daniel Jünger , Kevin Kristensen , Yunsong Wang , Xiangyao Yu , Bertil Schmidt

We describe an explicit construction of approximate Ginsparg-Wilson fermions for QCD. We use ingredients of perfect action origin, and further elements. The spectrum of the lattice Dirac operator reveals the quality of the approximation. We…

High Energy Physics - Lattice · Physics 2009-10-31 W. Bietenholz , N. Eicker , I. Hip , K. Schilling

We port Domain-Decomposed-alpha-AMG solver to the K computer. The system has 8 cores and 16 GB memory per node, of which theoretical peak is 128 GFlops (82,944 nodes in total). Its feature, as many as 256 registers per core and as large as…

High Energy Physics - Lattice · Physics 2019-04-02 Ken-Ichi Ishikawa , Issaku Kanamori

We present a parallel computing strategy for a hybridizable discontinuous Galerkin (HDG) nested geometric multigrid (GMG) solver. Parallel GMG solvers require a combination of coarse-grain and fine-grain parallelism to improve time to…

Numerical Analysis · Mathematics 2019-07-18 M. S. Fabien , M. G. Knepley , R. T. Mills , B. M. Riviere

The high degree of parallelism and relatively complicated synchronization mechanisms in GPUs make writing correct kernels difficult. Data races pose one such concurrency correctness challenge, and therefore, effective methods of detecting…

Programming Languages · Computer Science 2021-11-25 Sagnik Dey , Mayant Mukul , Parth Sharma , Swarnendu Biswas

To support the running of human-centric metaverse applications on mobile devices, Unmanned Aerial Vehicle (UAV)-assisted Wireless Powered Mobile Edge Computing (WPMEC) is promising to compensate for limited computational capabilities and…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-11-29 Xiaojie Wang , Jiameng Li , Zhaolong Ning , Qingyang Song , Lei Guo , Abbas Jamalipour

This paper presents the design, implementation, and performance analysis of a parallel and GPU-accelerated Poisson solver based on the Preconditioned Bi-Conjugate Gradient Stabilized (Bi-CGSTAB) method. The implementation utilizes the MPI…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-03-13 Luca Pennati , Måns I. Andersson , Klaus Steiniger , Rene Widera , Tapish Narwal , Michael Bussmann , Stefano Markidis

A simple minded approach to implement three discretizations of the Dirac operator (staggered, Wilson, Brillouin) on two architectures (KNL and core i7) is presented. The idea is to use a high-level compiler along with OpenMP parallelization…

High Energy Physics - Lattice · Physics 2018-11-21 Stephan Durr

Pauli Correlation Encoding (PCE) is as a qubit-efficient variational approach to combinatorial optimization problems. The method offers a polynomial reduction in qubit count and a super-polynomial suppression of barren plateaus. Here, we…

We explore the use of the Cell Broadband Engine (Cell/BE for short) for combinatorial optimization applications: we present a parallel version of a constraint-based local search algorithm that has been implemented on a multiprocessor…

Artificial Intelligence · Computer Science 2009-10-08 Salvator Abreu , Daniel Diaz , Philippe Codognet

Accelerating Human Action Recognition (HAR) efficiently for real-time surveillance and robotic systems on edge chips remains a challenging research field, given its high computational and memory requirements. This paper proposed an…

Computer Vision and Pattern Recognition · Computer Science 2023-11-08 Azzam Alhussain , Mingjie Lin

Immense interest in quantum computing has prompted development of electronic structure methods that are suitable for quantum hardware. However, the slow pace at which quantum hardware progresses, forces researchers to implement their ideas…

Quantum Physics · Physics 2025-02-26 Ilya G. Ryabinkin , Seyyed Mehdi Hosseini Jenab , Scott N. Genin
‹ Prev 1 8 9 10 Next ›