Related papers: Evaluation of disconnected contributions using GPU…
This work concerns the numerical simulation of the Vlasov-Poisson set of equations using semi- Lagrangian methods on Graphical Processing Units (GPU). To accomplish this goal, modifications to traditional methods had to be implemented.…
In this paper, we develop a multiscale finite element method for solving flows in fractured media. Our approach is based on Generalized Multiscale Finite Element Method (GMsFEM), where we represent the fracture effects on a coarse grid via…
We study a graph partitioning problem motivated by the simulation of the physical movement of multi-body systems on an atomistic level, where the forces are calculated from a quantum mechanical description of the electrons. Several advanced…
Recently proposed Graph Neural Networks (GNNs) for vertex clustering are trained with an unsupervised minimum cut objective, approximated by a Spectral Clustering (SC) relaxation. However, the SC relaxation is loose and, while it offers a…
Recent studies have shown that Binary Graph Neural Networks (GNNs) are promising for saving computations of GNNs through binarized tensors. Prior work, however, mainly focused on algorithm designs or training techniques, leaving it open to…
An implementation of the generalized time-dependent generator coordinated method (TD-GCM) is developed, that can be applied to the dynamics of small- and large-amplitude collective motion of atomic nuclei. Both the generator states and…
We present first results from our analysis of the most general quark-quark correlator of the nucleon, which can be parameterized in terms of so-called generalized transverse momentum dependent parton distributions. These results include the…
The predictions of the geometric collective model (GCM) for different sets of Hamiltonian parameter values are related by analytic scaling relations. For the quartic truncated form of the GCM -- which describes harmonic oscillator, rotor,…
This work proposes a GPU tensor core approach that encodes the arithmetic reduction of $n$ numbers as a set of chained $m \times m$ matrix multiply accumulate (MMA) operations executed in parallel by GPU tensor cores. The asymptotic running…
Functional graphical models explore dependence relationships of random processes. This is achieved through estimating the precision matrix of the coefficients from the Karhunen-Loeve expansion. This paper deals with the problem of…
By building on recent advances in the use of randomized trace estimation to drastically reduce the memory footprint of adjoint-state methods, we present and validate an imaging approach that can be executed exclusively on accelerators.…
An efficient solver for the three dimensional free-space Poisson equation is presented. The underlying numerical method is based on finite Fourier series approximation. While the error of all involved approximations can be fully controlled,…
Disaggregation maps parts of an AI workload to different types of GPUs, offering a path to utilize modern heterogeneous GPU clusters. However, existing solutions operate at a coarse granularity and are tightly coupled to specific model…
A recent Graph Neural Network (GNN) approach for learning to branch has been shown to successfully reduce the running time of branch-and-bound algorithms for Mixed Integer Linear Programming (MILP). While the GNN relies on a GPU for…
The calculation of the nucleon strangeness form factors from N_f=2+1 clover fermion lattice QCD is presented. Disconnected insertions are evaluated using the Z(4) stochastic method, along with unbiased subtractions from the hopping…
Graph Neural Networks (GNN) exhibit superior performance in graph representation learning, but their inference cost can be high, due to an aggregation operation that can require a memory fetch for a very large number of nodes. This…
The prediction of a dielectric breakdown in a high-voltage device is based on criteria that evaluate the electric field along field lines. Therefore it is necessary to efficiently compute the electric field at arbitrary points in space. A…
When training a Neural Network, it is optimized using the available training data with the hope that it generalizes well to new or unseen testing data. At the same absolute value, a flat minimum in the loss landscape is presumed to…
We reanalyze the experimental NMC data on the nonsinglet structure function $F_2^p-F_2^n$ and E866 data on the nucleon sea asymmetry $\bar{d}/\bar{u}$ using the truncated moments approach elaborated in our previous papers. With help of the…
The variational inclusion of spin-orbit coupling in self-consistent field (SCF) calculations requires a generalised two-component framework, which permits the single-determinant wave function to completely break spin symmetry. The…