English
Related papers

Related papers: Towards optimal parallel PM N-body codes: PMFAST

200 papers

We present two new algebraic multilevel hierarchical matrix algorithms to perform fast matrix-vector product (MVP) for $N$-body problems in $d$ dimensions, namely efficient $\mathcal{H}^2_{*}$ (fully nested algorithm, i.e., $\mathcal{H}^2$…

Numerical Analysis · Mathematics 2026-04-13 Ritesh Khan , Sivaram Ambikasaran

Recent increases in supercomputing power, driven by the multi-core revolution and accelerators such as the IBM Cell processor, graphics processing units (GPUs) and Intel's Many Integrated Core (MIC) technology have enabled kinetic…

We present a high-performance N-body code for self-gravitating collisional systems accelerated with the aid of a new SIMD instruction set extension of the x86 architecture: Advanced Vector eXtensions (AVX), an enhanced version of the…

Instrumentation and Methods for Astrophysics · Physics 2015-05-27 Ataru Tanikawa , Kohji Yoshikawa , Takashi Okamoto , Keigo Nitadori

We present the multi-GPU realization of the StePS (Stereographically Projected Cosmological Simulations) algorithm with MPI-OpenMP-CUDA hybrid parallelization and nearly ideal scale-out to multiple compute nodes. Our new zoom-in…

Cosmology and Nongalactic Astrophysics · Physics 2019-03-22 Gábor Rácz , István Szapudi , László Dobos , István Csabai , Alexander S. Szalay

This paper presents CUBEP3M, a publicly-available high performance cosmological N-body code and describes many utilities and extensions that have been added to the standard package. These include a memory-light runtime SO halo finder, a…

Cosmology and Nongalactic Astrophysics · Physics 2015-06-11 Joachim Harnois-Deraps , Ue-Li Pen , Ilian T. Iliev , Hugh Merz , J. D. Emberson , Vincent Desjacques

This article presents MuMFiM, an open source application for multiscale modeling of fibrous materials on massively parallel computers. MuMFiM uses two scales to represent fibrous materials such as biological network materials (extracellular…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-04 Jacob Merson , Catalin Picu , Mark S. Shephard

Application partitioning and code offloading are being researched extensively during the past few years. Several frameworks for code offloading have been proposed. However, fewer works attempted to address issues occurred with its…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-09-21 Nevin Vunka Jungum , Nawaz Mohamudally , Nimal Nissanke

N-body simulations are essential tools in physical cosmology to understand the large-scale structure (LSS) formation of the Universe. Large-scale simulations with high resolution are important for exploring the substructure of universe and…

Computational Physics · Physics 2020-10-22 Shenggan Cheng , Hao-Ran Yu , Derek Inman , Qiucheng Liao , Qiaoya Wu , James Lin

In this article, we present Defmod, an open source, fully unstructured, two or three dimensional, parallel finite element code for modeling crustal deformation over time scales ranging from milliseconds to thousands of years. Unlike…

Geophysics · Physics 2015-12-31 S. Tabrez Ali

We present the basic idea, implementation, measured performance and performance model of FDPS (Framework for developing particle simulators). FDPS is an application-development framework which helps the researchers to develop particle-based…

Instrumentation and Methods for Astrophysics · Physics 2016-06-15 Masaki Iwasawa , Ataru Tanikawa , Natsuki Hosono , Keigo Nitadori , Takayuki Muranushi , Junichiro Makino

A new, momentum preserving fast Poisson solver for N-body systems sharing effective O(N) computational complexity, has been recently developed by Dehnen (2000, 2002). We have implemented the proposed algorithms in a Fortran-90 code, and…

Astrophysics · Physics 2007-05-23 P. Londrillo , C. Nipoti , L. Ciotti

Laser plasma instabilities (LPIs) have significant influences on the laser energy deposition efficiency, hot electron generation, and uniformity of irradiation in inertial confined fusion (ICF). In contrast to theoretical analysis of linear…

Plasma Physics · Physics 2024-04-23 Hanghang Ma , Liwei Tan , Suming Weng , Wenjun Ying , Zhengming Sheng , Jie Zhang

We present a novel convex formulation that weakly couples the Material Point Method (MPM) with rigid body dynamics through frictional contact, optimized for efficient GPU parallelization. Our approach features an asynchronous time-splitting…

Robotics · Computer Science 2025-07-08 Chang Yu , Wenxin Du , Zeshun Zong , Alejandro Castro , Chenfanfu Jiang , Xuchen Han

Balanced hypergraph partitioning is an NP-hard problem with many applications, e.g., optimizing communication in distributed data placement problems. The goal is to place all nodes across $k$ different blocks of bounded size, such that…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-04-03 Lars Gottesbüren , Tobias Heuer , Nikolai Maas , Peter Sanders , Sebastian Schlag

Over the last two decades, several fast, robust, and high-order accurate methods have been developed for solving the Poisson equation in complicated geometry using potential theory. In this approach, rather than discretizing the partial…

Numerical Analysis · Mathematics 2024-09-19 Fredrik Fryklund , Leslie Greengard , Shidong Jiang , Samuel Potter

Today's exponentially increasing data volumes and the high cost of storage make compression essential for the Big Data industry. Although research has concentrated on efficient compression, fast decompression is critical for analytics…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-06-03 Evangelia Sitaridi , Rene Mueller , Tim Kaldewey , Guy Lohman , Kenneth Ross

A novel splitting scheme to solve parametric multiconvex programs is presented. It consists of a fixed number of proximal alternating minimisations and a dual update per time step, which makes it attractive in a real-time Nonlinear Model…

Optimization and Control · Mathematics 2014-07-22 Jean-Hubert Hours , Colin N. Jones

In this paper we present and evaluate a parallel algorithm for solving a minimum spanning tree (MST) problem for supercomputers with distributed memory. The algorithm relies on the relaxation of the message processing order requirement for…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-10-18 Artem Mazeev , Alexander Semenov , Alexey Simonov

We present a work-efficient parallel level-synchronous Breadth First Search (BFS) algorithm for shared-memory architectures which achieves the theoretical lower bound on parallel running time. The optimality holds regardless of the shape of…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-09-20 Jesmin Jahan Tithi , Yonatan Fogel , Rezaul Chowdhury

This paper applies the N-block PCPM algorithm to solve multi-scale multi-stage stochastic programs, with the application to electricity capacity expansion models. Numerical results show that the proposed simplified N-block PCPM algorithm,…

Optimization and Control · Mathematics 2021-03-29 Run Chen , Andrew L. Liu