English
Related papers

Related papers: CATAPULT: A CUDA-Accelerated Timestepper for Alpha…

200 papers

We present recent developments in the parallelization scheme of ECHO-3DHPC, an efficient astrophysical code used in the modelling of relativistic plasmas. With the help of the Intel Software Development Tools, like Fortran compiler and…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-10-11 Matteo Bugli , Luigi Iapichino , Fabio Baruffa

In this work, we optimize speculative sampling for parallel hardware accelerators to improve sampling speed. We notice that substantial portions of the intermediate matrices necessary for speculative sampling can be computed concurrently.…

Machine Learning · Computer Science 2024-10-04 Dominik Wagner , Seanie Lee , Ilja Baumann , Philipp Seeberger , Korbinian Riedhammer , Tobias Bocklet

A modern graphics processing unit (GPU) is able to perform massively parallel scientific computations at low cost. We extend our implementation of the checkerboard algorithm for the two dimensional Ising model [T. Preis et al., J. Comp.…

Computational Physics · Physics 2010-07-22 Benjamin Block , Peter Virnau , Tobias Preis

In the advent of new large galaxy surveys, which will produce enormous datasets with hundreds of millions of objects, new computational techniques are necessary in order to extract from them any two-point statistic, the computational time…

Instrumentation and Methods for Astrophysics · Physics 2013-06-21 David Alonso

Scan (or prefix sum) is a fundamental and widely used primitive in parallel computing. In this paper, we present LightScan, a faster parallel scan primitive for CUDA-enabled GPUs, which investigates a hybrid model combining intra-block…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-04-19 Yongchao Liu , Srinivas Aluru

The Center for Exascale Monte Carlo Neutron Transport is developing Monte Carlo / Dynamic Code (MC/DC) as a portable Monte Carlo neutron transport package for rapid numerical methods exploration on CPU- and GPU-based high-performance…

Computational Physics · Physics 2025-05-30 Joanna Piper Morgan , Braxton Cuneo , Ilham Variansyah , Kyle E. Niemeyer

A finite-difference Micromagnetic solver is presented utilizing the C++ Accelerated Massive Parallelism (C++ AMP). The high speed performance of a single Graphics Processing Unit (GPU) is demonstrated compared to a typical CPU-based solver.…

Computational Engineering, Finance, and Science · Computer Science 2014-07-07 Ru Zhu

The study deals with the parallelization of 2D and 3D finite element based Navier-Stokes codes using direct solvers. Development of sparse direct solvers using multifrontal solvers has significantly reduced the computational time of direct…

Mathematical Software · Computer Science 2009-10-13 Mandhapati P. Raju

The numerical integration of stochastic trajectories to estimate the time to pass a threshold is an interesting physical quantity, for instance in Josephson junctions and atomic force microscopy, where the full trajectory is not accessible.…

Computational Physics · Physics 2018-02-15 Vincenzo Pierro , Luigi Troiano , Elena Mejuto , Giovannni Filatrella

Adiabatic quantum computers, such as the quantum annealers commercialized by D-Wave Systems Inc., are routinely used to tackle combinatorial optimization problems. In this article, we show how to exploit them to accelerate equilibrium…

Disordered Systems and Neural Networks · Physics 2023-07-12 Giuseppe Scriva , Emanuele Costa , Benjamin McNaughton , Sebastiano Pilati

Linear Programs (LPs) appear in a large number of applications and offloading them to a GPU is viable to gain performance. Existing work on offloading and solving an LP on a GPU suggests that there is performance gain generally on large…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-02-26 Amit Gurung , Rajarshi Ray

This paper introduces cuVegas, a CUDA-based implementation of the Vegas Enhanced Algorithm (VEGAS+), optimized for multi-dimensional integration in GPU environments. The VEGAS+ algorithm is an advanced form of Monte Carlo integration,…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-08-20 Emiliano Tolotti , Anas Jnini , Flavio Vella , Roberto Passerone

Clawpack is a library for solving nonlinear hyperbolic partial differential equations using high-resolution finite volume methods based on Riemann solvers and limiters. It supports Adaptive Mesh Refinement (AMR), which is essential in…

Mathematical Software · Computer Science 2018-08-09 Xinsheng Qin , Randall J. LeVeque , Michael R. Motley

Modern graphics processing units (GPUs) provide impressive computing resources, which can be accessed conveniently through the CUDA programming interface. We describe how GPUs can be used to considerably speed up molecular dynamics (MD)…

Computational Physics · Physics 2011-04-08 Peter H. Colberg , Felix Höfling

The High Energy Physics (HEP) experiments, such as those at the Large Hadron Collider (LHC), traditionally consume large amounts of CPU cycles for detector simulations and data analysis, but rarely use compute accelerators such as GPUs. As…

High Energy Physics - Experiment · Physics 2022-03-17 Zhihua Dong , Heather Gray , Charles Leggett , Meifeng Lin , Vincent R. Pascuzzi , Kwangmin Yu

A spectral fitter based on the graphics processor unit (GPU) has been developed for Borexino solar neutrino analysis. It is able to shorten the fitting time to a superior level compared to the CPU fitting procedure. In Borexino solar…

Data Analysis, Statistics and Probability · Physics 2020-01-22 X. F. Ding , M. Agostini , K. Altenmuller , S. Appel , V. Atroshchenko , Z. Bagdasarian , D. Basilico , G. Bellini , J. Benziger , D. Bick , G. Bonfini , D. Bravo , B. Caccianiga , F. Calaprice , A. Caminata , S. Caprioli , M. Carlini , P. Cavalcante , A. Chepurnov , K. Choi , L. Collica , D. D'Angelo , S. Davini , A. Derbin , A. Di Ludovico , L. Di Noto , I. Drachnev , K. Fomenko , A. Formozov , D. Franco , F. Froborg , F. Gabriele , C. Galbiati , C. Ghiano , M. Giammarchi , A. Goretti , M. Gromov , D. Guffanti , C. Hagner , T. Houdy , E. Hungerford , Aldo Ianni , Andrea Ianni , A. Jany , D. Jeschke , V. Kobychev , D. Korablev , G. Korga , D. Kryn , M. Laubenstein , E. Litvinovich , F. Lombardi , P. Lombardi , L. Ludhova , G. Lukyanchenko , L. Lukyanchenko , I. Machulin , G. Manuzio , S. Marcocci , J. Martyn , E. Meroni , M. Meyer , L. Miramonti , M. Misiaszek , V. Muratova , B. Neumair , L. Oberauer , B. Opitz , V. Orekhov , F. Ortica , M. Pallavicini , L. Papp , O. Penek , N. Pilipenko , A. Pocar , A. Porcelli , G. Ranucci , A. Razeto , A. Re , M. Redchuk , A. Romani , R. Roncin , N. Rossi , S. Schonert , D. Semenov , M. Skorokhvatov , O. Smirnov , A. Sotnikov , L. F. F. Stokes , Y. Suvorov , R. Tartaglia , G. Testera , J. Thurn , M. Toropova , E. Unzhakov , A. Vishneva , R. B. Vogelaar , F. von Feilitzsch , H. Wang , S. Weinz , M. Wojcik , M. Wurm , Z. Yokley , O. Zaimidoroga , S. Zavatarelli , K. Zuber , G. Zuzel

We describe the GPU implementation of shifted or multimass iterative solvers for sparse linear systems of the sort encountered in lattice gauge theory. We provide a generic tool that can be used by those without GPU programming experience…

High Energy Physics - Lattice · Physics 2011-02-16 Richard Galvez , Greg van Anders

We present a multi-purpose genetic algorithm, designed and implemented with GPGPU / CUDA parallel computing technology. The model was derived from our CPU serial implementation, named GAME (Genetic Algorithm Model Experiment). It was…

Instrumentation and Methods for Astrophysics · Physics 2015-06-15 Stefano Cavuoti , Mauro Garofalo , Massimo Brescia , Maurizio Paolillo , Antonio Pescape' , Giuseppe Longo , Giorgio Ventre

Numerical integration of stochastic differential equations is commonly used in many branches of science. In this paper we present how to accelerate this kind of numerical calculations with popular NVIDIA Graphics Processing Units using the…

Computational Physics · Physics 2011-05-31 M. Januszewski , M. Kostur

We introduce XtraPuLP, a new distributed-memory graph partitioner designed to process trillion-edge graphs. XtraPuLP is based on the scalable label propagation community detection technique, which has been demonstrated as a viable means to…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-10-25 George M Slota , Sivasankaran Rajamanickam , Karen Devine , Kamesh Madduri