中文
相关论文

相关论文: Domain Decomposition method on GPU cluster

200 篇论文

We consider one-level additive Schwarz domain decomposition preconditioners for the Helmholtz equation with variable coefficients (modelling wave propagation in heterogeneous media), subject to boundary conditions that include wave…

数值分析 · 数学 2020-10-06 Shihua Gong , Ivan G. Graham , Euan A. Spence

The present work describes the development of heterogeneous GPGPU implicit CFD coupled solvers, encompassing both density- and pressure- based approaches. In this setup, the assembled linear matrix is offloaded onto multiple GPUs using…

分布式、并行与集群计算 · 计算机科学 2024-03-14 Stefano Oliani , Ettore Fadiga , Ivan Spisso , Luigi Capone , Federico Piscaglia

Barren plateaus present a major challenge in the training of variational quantum algorithms (VQAs), particularly for large-scale discretizations of nonlinear partial differential equations. In this work, we introduce a domain decomposition…

数值分析 · 数学 2026-03-26 Laila S. Busaleh , Jeonghyeuk Kwon , Orlane Zang , Muhammad Hassan , Yvon Maday

When designing a power or CPU constrained device where a four-axis robotic arm is required and access to the Robot Operating System (ROS) is not an option, finding an efficient state space controller for a four-axis arm can be an obstacle.…

机器人学 · 计算机科学 2023-12-25 Alistair Keiller

Sparse direct linear solvers are at the computational core of domain decomposition preconditioners and therefore have a strong impact on their performance. In this paper, we consider the Fast and Robust Overlapping Schwarz (FROSch) solver…

The Variational Quantum Linear Solver (VQLS), a hybrid quantum-classical algorithm for solving linear systems, faces a practical scalability bottleneck: the Linear Combination of Unitaries (LCU) decomposition requires O(L^2) circuit…

Improving time-to-solution in molecular dynamics simulations often requires strong scaling due to fixed-sized problems. GROMACS is highly latency-sensitive, with peak iteration rates in the sub-millisecond, making scalability on…

分布式、并行与集群计算 · 计算机科学 2025-09-29 Mahesh Doijade , Andrey Alekseenko , Ania Brown , Alan Gray , Szilárd Páll

We introduce a class of efficient multiple right-hand side multigrid algorithm for domain wall fermions. The simultaneous solution for a modest number of right hand sides concurrently allows for a significant reduction in the time spent…

高能物理 - 格点 · 物理学 2024-09-09 Peter A Boyle

We present an efficient, robust and fully GPU-accelerated aggregation-based algebraic multigrid preconditioning technique for the solution of large sparse linear systems. These linear systems arise from the discretization of elliptic PDEs.…

数值分析 · 数学 2014-03-10 Rajesh Gandham , Ken Esler , Yongpeng Zhang

We consider one-level additive Schwarz preconditioners for a family of Helmholtz problems with absorption and increasing wavenumber $k$. These problems are discretized using the Galerkin method with nodal conforming finite elements of any…

数值分析 · 数学 2020-05-20 I. G. Graham , E. A. Spence , J. Zou

We propose a preconditioner to accelerate the convergence of the GMRES iterative method for solving the system of linear equations obtained from discretize-then-optimize approach applied to optimal control problems constrained by a partial…

数值分析 · 数学 2019-11-15 Hamid Mirchi , Davod Khojasteh Salkuyeh

We present a matrix-free GPU multigrid preconditioner with algebraically consistent coarsening for solving Poisson equations on adaptive octree grids with irregular domains. Within uniform-resolution regions, the coarsening satisfies the…

数值分析 · 数学 2026-04-22 Mengdi Wang , Yuchen Sun , Bo Zhu

We present a GPU-accelerated version of the real-space SPARC electronic structure code for performing hybrid functional calculations in generalized Kohn-Sham density functional theory. In particular, we develop a batch variant of the…

计算物理 · 物理学 2025-01-29 Xin Jing , Abhiraj Sharma , John E. Pask , Phanish Suryanarayana

We present a parallel implementation of a direct solver for the Poisson's equation on extreme-scale supercomputers with accelerators. We introduce a chunked-pencil decomposition as the domain-decomposition strategy to distribute work among…

计算物理 · 物理学 2020-07-15 Jaber J. Hasbestan , Inanc Senocak

Recommender systems rely heavily on increasing computation resources to improve their business goal. By deploying computation-intensive models and algorithms, these systems are able to inference user interests and exhibit certain ads or…

系统与控制 · 电气工程与系统科学 2021-03-04 Xun Yang , Yunli Wang , Cheng Chen , Qing Tan , Chuan Yu , Jian Xu , Xiaoqiang Zhu

Lattice QCD simulations are computationally expensive, with the solution of the Dirac equation being the major computational bottleneck of many calculations. We introduce a novel gauge-equivariant neural-network architecture for…

高能物理 - 格点 · 物理学 2026-04-23 Simon Pfahler , Daniel Knüttel , Christoph Lehner , Tilo Wettig

We study a variant of the Schwarz-preconditioned HMC algorithm. In contrast to the original proposal of L\"uscher, we apply the domain decomposition in one lattice direction only. This is sufficient to reduce the condition number of the…

高能物理 - 格点 · 物理学 2008-11-26 Martin Hasenbusch

The convergence rate of domain decomposition methods (DDMs) strongly depends on the transmission condition at the interfaces between subdomains. Thus, an important aspect in improving the efficiency of such solvers is careful design of…

数值分析 · 数学 2025-05-19 Niall Bootland , Sahar Borzooei , Victorita Dolean , Pierre-Henri Tournier

In order to satisfy timing constraints, modern real-time applications require massively parallel accelerators such as General Purpose Graphic Processing Units (GPGPUs). Generation after generation, the number of computing clusters made…

分布式、并行与集群计算 · 计算机科学 2021-05-24 Houssam-Eddine Zahaf , Ignacio Sanudo Olmedo , Jayati Singh , Nicola Capodieci , Sebastien Faucou

Basic Linear Algebra Subprograms (BLAS) play key role in high performance and scientific computing applications. Experimentally, yesteryear multicore and General Purpose Graphics Processing Units (GPGPUs) are capable of achieving up to 15…

硬件体系结构 · 计算机科学 2016-11-29 Farhad Merchant , Tarun Vatwani , Anupam Chattopadhyay , Soumyendu Raha , S K Nandy , Ranjani Narayan