Related papers: GPU implementation of a Landau gauge fixing algori…
Computational fluid dynamics and fluid-structure interaction simulations involving moving and deforming bodies is extremely hard. In this work, we present a graphical processing unit (GPU) optimized implementation of the sharp-interface…
We propose a method which allows the generalization of the Landau lattice gauge-fixing procedure to generic covariant gauges. We report preliminary numerical results showing how the procedure works for $SU(2)$ and $SU(3)$. We also report…
In this paper we test an approximate method that is often used in lattice studies of the Landau gauge three-gluon vertex. The approximation consists in describing the lattice correlator with tensor bases from the continuum theory. With the…
A fast algorithm for the approximation of a low rank LU decomposition is presented. In order to achieve a low complexity, the algorithm uses sparse random projections combined with FFT-based random projections. The asymptotic approximation…
Numerical modeling of nematic liquid crystals using the tensorial Landau-de Gennes (LdG) theory provides detailed insights into the structure and energetics of the enormous variety of possible topological defect configurations that may…
We evaluate the three-gluon vertex with one vanishing external momentum within the Curci-Ferrari (CF) model at two-loop order and compare our results to Landau-gauge lattice simulations of the same vertex function for the SU(2) and SU(3)…
We discuss a new lattice implementation of the linear covariant gauge, recently introduced in [1]. In particular, we present details of the numerical procedure for fixing the gauge. We also report on preliminary results for the transverse…
LU factorization for sparse matrices is the most important computing step for many engineering and scientific computing problems such as circuit simulation. But parallelizing LU factorization with the Graphic Processing Units (GPU) still…
Parallel computing can offer an enormous advantage regarding the performance for very large applications in almost any field: scientific computing, computer vision, databases, data mining, and economics. GPUs are high performance many-core…
We compare two Landau gauge fixing methods, aiming to find the global maximum of the gauge fixing functional. Moreover, a systematic effect of Gribov copies in the gluon and ghost propagators computed in Landau gauge is presented and…
Lattice gauge fixing is required to compute gauge-variant quantities, for example those used in RI-MOM renormalization schemes or as objects of comparison for model calculations. Recently, gauge-variant quantities have also been found to be…
This paper presents a Graphics Processing Units (GPUs) acceleration method of an iterative scheme for gas-kinetic model equations. Unlike the previous GPU parallelization of explicit kinetic schemes, this work features a fast converging…
Here we present an implementation of Primal-Dual Affine scaling method to solve linear optimization problem on GPU based systems. Strategies to convert the system generated by complementary slackness theorem into a symmetric system are…
The acceleration of sparse matrix computations on modern many-core processors, such as the graphics processing units (GPUs), has been recognized and studied over a decade. Significant performance enhancements have been achieved for many…
We present a new adaptive parallel algorithm for the challenging problem of multi-dimensional numerical integration on massively parallel architectures. Adaptive algorithms have demonstrated the best performance, but efficient many-core…
Lattice spin models are useful for studying critical phenomena and allow the extraction of equilibrium and dynamical properties. Simulations of such systems are usually based on Monte Carlo (MC) techniques, and the main difficulty is often…
We present a new method of gauge fixing to standard lattice Landau gauge, Max Re Tr $\sum_{\mu,x}U_{\mu,x}$, in which the link configuration is recursively smeared; these smeared links are then gauge fixed by standard extremization. The…
In this work, we introduce a new hardware architecture for decoding correlated errors in quantum LDPC codes. The decoder is based on message passing and exploits the structure of the detector error model obtained through the recently…
An existing hybrid MPI-OpenMP scheme is augmented with a CUDA-based fine grain parallelization approach for multidimensional distributed Fourier transforms, in a well-characterized pseudospectral fluid turbulence code. Basics of the hybrid…
We present LBcuda, a GPU accelerated version of LBsoft, our open-source MPI-based software for the simulation of multi-component colloidal flows. We describe the design principles, the optimization and the resulting performance as compared…