Related papers: THOR: a GPU-accelerated and MPI-parallel radiative…
Making general particle transport simulation for high-energy physics (HEP) single-instruction-multiple-thread (SIMT) friendly, to take advantage of accelerator hardware, is an important alternative for boosting the throughput of simulation…
We present a new library for parallel distributed Fast Fourier Transforms (FFT). The importance of FFT in science and engineering and the advances in high performance computing necessitate further improvements. AccFFT extends existing FFT…
To study the atomic, molecular and ionized emission of Giant Molecular Clouds (GMCs), we have initiated a Large Program with the VLA: 'THOR - The HI, OH, Recombination Line survey of the Milky Way'. We map the 21cm HI line, 4 OH lines, 19…
We present the extension of the differentiable hydrodynamics code, diffhydro, enabling scalable PDE-constrained inference and integrated hybrid physics-ML models for a wide range of astrophysical applications. New physics additions include…
We present a new very fast tree-code which runs on massively parallel Graphical Processing Units (GPU) with NVIDIA CUDA architecture. The tree-construction and calculation of multipole moments is carried out on the host CPU, while the force…
The development of fast numerical methods for multilevel radiative transfer (RT) applications often leads to important breakthroughs in astrophysics, because they allow the investigation of problems that could not be properly tackled using…
OCTO-TIGER is an astrophysics code to simulate the evolution of self-gravitating and rotat-ing systems of arbitrary geometry based on the fast multipole method, using adaptive mesh refinement. OCTO-TIGER is currently optimised to simulate…
We present a Semi-Analytical Line Transfer model, SALT, to study the absorption and re-emission line profiles from expanding galactic envelopes. The envelopes are described as a superposition of shells with density and velocity varying with…
We present MARUT, a scalable multi-GPU computational fluid dynamics (CFD) framework designed for high-fidelity simulations of compressible flows spanning subsonic to hypersonic regimes, including chemically reacting nonequilibrium flows…
We introduce the CUDA Tensor Transpose (cuTT) library that implements high-performance tensor transposes for NVIDIA GPUs with Kepler and above architectures. cuTT achieves high performance by (a) utilizing two GPU-optimized transpose…
The increased bandwidth coupled with the large numbers of antennas of several new radio telescope arrays has resulted in an exponential increase in the amount of data that needs to be recorded and processed. In many cases, it is necessary…
We present a novel numerical implementation of radiative transfer in the cosmological smoothed particle hydrodynamics (SPH) simulation code {\small GADGET}. It is based on a fast, robust and photon-conserving integration scheme where the…
We present MGPU, a C++ programming library targeted at single-node multi-GPU systems. Such systems combine disproportionate floating point performance with high data locality and are thus well suited to implement real-time algorithms. We…
Non-LTE radiative transfer is a key tool for modern astrophysics: it is the means by which many key synthetic observables are produced, thus connecting simulations and observations. Radiative transfer models also inform our understanding of…
We present a highly parallel implementation of the cross-correlation of time-series data using graphics processing units (GPUs), which is scalable to hundreds of independent inputs and suitable for the processing of signals from "Large-N"…
Solar flares involve complex processes that are coupled and span a wide range of temporal, spatial, and energy scales. Modeling such processes self-consistently has been a challenge in the past. Here we present results from simulations that…
In this paper we present CRASH_alpha, the first radiative transfer code for cosmological application that follows the parallel propagation of Ly_alpha and ionizing photons. CRASH_alpha is a version of the continuum radiative transfer code…
Modern compute nodes in high-performance computing provide a tremendous level of parallelism and processing power. However, as arithmetic performance has been observed to increase at a faster rate relative to memory and network bandwidths,…
Ly$\alpha$ intensity mapping is emerging as a new probe of faint galaxies consisting the cosmic web that elude traditional surveys. However, the resonant nature of Ly$\alpha$ radiative transfer complicates the interpretation of observed…
To assess how future progress in gravitational microlensing computation at high optical depth will rely on both hardware and software solutions, we compare a direct inverse ray-shooting code implemented on a graphics processing unit (GPU)…