Related papers: Parallel implementation of the Density Matrix Reno…
Modern datacenters increasingly rely on low-power, single-slot inference accelerators to balance performance, energy efficiency, and rack density constraints. The NVIDIA T4 GPU has become widely deployed due to strong performance per watt…
We present a method for computing resonant inelastic x-ray scattering (RIXS) spectra in one-dimensional systems using the density matrix renormalization group (DMRG) method. By using DMRG to address the problem, we shift the computational…
We study the application of the density matrix renormalization group (DMRG) to systems with one-dimensional acoustic phonons. We show how the use of a local oscillator basis circumvents the difficulties with the long-range interactions…
The numerical renormalization group (NRG) is rephrased as a variational method with the cost function given by the sum of all the energies of the effective low-energy Hamiltonian. This allows to systematically improve the spectrum obtained…
A new approach to large-scale nuclear structure calculations, based on the Density Matrix Renormalization Group (DMRG), is described. The method is tested in the context of a problem involving many identical nucleons constrained to move in…
We have developed an efficient method for performing density matrix renormalization group (DMRG) simulations of the SU(N) Fermi-Hubbard chain with open boundary conditions, fully leveraging the SU(N) symmetry of the problem. This method…
Compared to ground state electronic structure optimizations, accurate simulations of molecular real-time electron dynamics are usually much more difficult to perform. To simulate electron dynamics, the time-dependent density matrix…
We propose a GPU-accelerated distributed optimization algorithm for controlling multi-phase optimal power flow in active distribution systems with dynamically changing topologies. To handle varying network configurations and enable…
A density-matrix renormalization group (DMRG) method for highly anisotropic two-dimensional systems is presented. The method consists in applying the usual DMRG in two steps. In the first step, a pure one dimensional calculation along the…
We propose a novel many-body framework combining the density matrix renormalization group (DMRG) with the valence-space (VS) formulation of the in-medium similarity renormalization group. This hybrid scheme admits for favorable…
We study the one-dimensional $S=1/2$ Heisenberg model with a uniform and a staggered magnetic fields, using the dynamical density-matrix renormalization group (DDMRG) technique. The DDMRG enables us to investigate the dynamical properties…
We present a GPU implementation of LAMMPS, a widely-used parallel molecular dynamics (MD) software package, and show 5x to 13x single node speedups versus the CPU-only version of LAMMPS. This new CUDA package for LAMMPS also enables…
Implicit methods and GPU parallelization are two distinct yet powerful strategies for accelerating high-order CFD algorithms. However, few studies have successfully integrated both approaches within high-speed flow solvers. The core…
We show that numerical computations based on tensor renormalization group (TRG) methods can be significantly accelerated with PyTorch on graphics processing units (GPUs) by leveraging NVIDIA's Compute Unified Device Architecture (CUDA). We…
We propose a new hybrid topology optimization algorithm based on multigrid approach that combines the parallelization strategy of CPU using OpenMP and heavily multithreading capabilities of modern Graphics Processing Units (GPU). In…
The exponential growth in data has intensified the demand for computational power to train large-scale deep learning models. However, the rapid growth in model size and complexity raises concerns about equal and fair access to computational…
Disaggregation maps parts of an AI workload to different types of GPUs, offering a path to utilize modern heterogeneous GPU clusters. However, existing solutions operate at a coarse granularity and are tightly coupled to specific model…
In this work, we consider the reformulation of hierarchical ($\mathcal{H}$) matrix algorithms for many-core processors with a model implementation on graphics processing units (GPUs). $\mathcal{H}$ matrices approximate specific dense…
A brief pedagogical overview of recent advances in tensor network state methods are presented that have the potential to broaden their scope of application radically for strongly correlated molecular systems. These include global fermionic…
Infinite projected entangled-pair states (iPEPS) provide a powerful tool for studying strongly correlated systems directly in the thermodynamic limit. A core component of the algorithm is the approximate contraction of the iPEPS, where the…