Related papers: Accelerating CCSD(T) on Graphical Processing Units…
The importance of post-CCSD(T) corrections as high as CCSDTQ56 for ground-state spectroscopic constants ($D_e$, $\omega_e$, $\omega_ex_e$, and $\alpha_e$) has been surveyed for a sample of two dozen mostly heavy-atom diatomics spanning a…
We develop a quartic-scaling implementation of coupled-cluster singles and doubles based on low-rank tensor hypercontraction (THC) factorizations of both the electron repulsion integrals (ERIs) and the doubles amplitudes. This extends our…
Correlation Power Analysis (CPA) is a type of power analysis based side channel attack that can be used to derive the secret key of encryption algorithms including DES (Data Encryption Standard) and AES (Advanced Encryption Standard). A…
The numerical integration of stochastic trajectories to estimate the time to pass a threshold is an interesting physical quantity, for instance in Josephson junctions and atomic force microscopy, where the full trajectory is not accessible.…
We present a cost-reduced approach for the distinguishable cluster approximation to coupled cluster with singles, doubles and iterative triples (DC-CCSDT) based on a tensor decomposition of the triples amplitudes. The triples amplitudes and…
We introduce a unitary coupled-cluster (UCC) ansatz termed $k$-UpCCGSD that is based on a family of sparse generalized doubles (D) operators which provides an affordable and systematically improvable unitary coupled-cluster wavefunction…
A number of iterative and perturbative approximations to the full equation-of-motion coupled cluster method with single, double, and triple excitations (EOM-CCSDT) are evaluated in the context of calculating the K-edge core-excitation and…
The NVIDIA Volta GPU microarchitecture introduces a specialized unit, called "Tensor Core" that performs one matrix-multiply-and-accumulate on 4x4 matrices per clock cycle. The NVIDIA Tesla V100 accelerator, featuring the Volta…
Radiation Treatment Planning (RTP) is the process of planning the appropriate external beam radiotherapy to combat cancer in human patients. RTP is a complex and compute-intensive task, which often takes a long time (several hours) to…
Quantum chemistry calculations of large, strongly correlated systems are typically limited by the computation cost that scales exponentially with the size of the system. Quantum algorithms, designed specifically for quantum computers, can…
Earth system models (ESM) demand significant hardware resources and energy consumption to solve atmospheric chemistry processes. Recent studies have shown improved performance from running these models on GPU accelerators. Nonetheless,…
Triangle counting (TC) is a fundamental problem in graph analysis and has found numerous applications, which motivates many TC acceleration solutions in the traditional computing platforms like GPU and FPGA. However, these approaches suffer…
We optimize matrix-product state-based algorithms for simulating quantum circuits with finite fidelity, specifically the time-evolving block decimation (TEBD) and the density-matrix renormalization group (DMRG) algorithms, by exploiting the…
Driven by deep learning, there has been a surge of specialized processors for matrix multiplication, referred to as TensorCore Units (TCUs). These TCUs are capable of performing matrix multiplications on small matrices (usually 4x4 or…
The increasing scale and wealth of inter-connected data, such as those accrued by social network applications, demand the design of new techniques and platforms to efficiently derive actionable knowledge from large-scale graphs. However,…
Sparse Matricized Tensor Times Khatri-Rao Product (spMTTKRP) is the bottleneck kernel of sparse tensor decomposition. In this work, we propose a GPU-based algorithm design to address the key challenges in accelerating spMTTKRP computation,…
We have developed a gravity solver based on combining the well developed Particle-Mesh (PM) method and TREE methods. It is designed for and has been implemented on parallel computer architectures. The new code can deal with tens of millions…
As transistor counts in a single chip exceed tens of billions, the complexity of RTL-level simulation and verification has grown exponentially, often extending simulation campaigns to several months. In industry practice, RTL simulation is…
Molecular dynamics (MD) simulation is a powerful computational tool to study the behavior of macromolecular systems. But many simulations of this field are limited in spatial or temporal scale by the available computational resource. In…
In this paper, we present a parallel computing method for the coupled-cluster singles and doubles (CCSD) in periodic systems. The CCSD in periodic systems solves simultaneous equations for single-excitation and double-excitation amplitudes.…