Related papers: Portably parallel construction of a CI wave functi…
Dynamic resource management is an increasingly important capability of High Performance Computing systems, as it enables jobs to adjust their resource allocation at runtime. This capability can reduce workload makespan, substantially…
Product distribution matching (PDM) is proposed to generate target distributions over large alphabets by combining the output of several parallel distribution matchers (DMs) with smaller output alphabets. The parallel architecture of PDM…
The combinatorial scaling of configuration interaction (CI) has long restricted its applicability to only the simplest molecular systems. Here, we report the first numerically exact CI calculation exceeding one quadrillion ($10^{15}$)…
The exponential scaling of complete active space (CAS) and full configuration interaction (CI) calculations limits the ability of quantum chemists to simulate the electronic structures of strongly correlated systems. Herein, we present…
We present a generative framework for inverse design of five-layer transmissive Huygens' metasurfaces (HMSs), addressing a longstanding challenge in achieving full-phase, high-efficiency unit cell designs with minimal full-wave simulations.…
Matrix Product State (MPS) is a versatile tensor network representation widely applied in quantum physics, quantum chemistry, and machine learning, etc. MPS sampling serves as a critical fundamental operation in these fields. As the…
The integration of quantum chemical methods with high-performance computing is indispensable for handling large systems with modest accuracy or even small systems but with high accuracy. Continuing with the unified implementation of…
Despite the increasing adoption of Field-Programmable Gate Arrays (FPGAs) in compute clouds, there remains a significant gap in programming tools and abstractions which can leverage network-connected, cloud-scale, multi-die FPGAs to…
To fully unlock the benefits of multiple-input multiple-output (MIMO) networks, downlink channel state information (CSI) is required at the base station (BS). In frequency division duplex (FDD) systems, the CSI is acquired through a…
Load balancing is a widely accepted technique for performance optimization of scientific applications on parallel architectures. Indeed, balanced applications do not waste processor cycles on waiting at points of synchronization and data…
Autonomous mobile robots (AMRs), used for search-and-rescue and remote exploration, require fast and robust planning and control schemes. Sampling-based approaches for Model Predictive Control, especially approaches based on the Model…
Parallel input performance issues are often neglected in large scale parallel applications in Computational Science and Engineering. Traditionally, there has been less focus on input performance because either input sizes are small (as in…
The Graspg program package is an extension of Grasp2018 [Comput. Phys. Commun. 237 (2019) 184-187] based on configuration state function generators (CSFGs). The generators keep spin-angular integrations at a minimum and reduce substantially…
The configuration interaction (CI) is a versatile wavefunction theory for interacting fermions but it involves an extremely long CI series. Using a symmetric tensor decomposition (STD) method, we convert the CI series into a compact and…
Accurate calculations of strongly correlated materials remain a formidable challenge in condensed matter physics, particularly due to the computational demand of conventional methods. This paper presents an efficient solver for dynamical…
The generalization of matrix product states (MPS) to continuous systems, as proposed in the breakthrough paper [F. Verstraete, J.I. Cirac, Phys. Rev. Lett. 104, 190405(2010)], provides a powerful variational ansatz for the ground state of…
Selected configuration interaction (sCI) methods including second-order perturbative corrections provide near full CI (FCI) quality energies with only a small fraction of the determinants of the FCI space. Here, we introduce both a…
We present a new wavefunction ansatz that combines the strengths of spin projection with the language of matrix product states (MPS) and matrix product operators (MPO) as used in the density matrix renormalization group (DMRG).…
In this paper, we present a parallel algorithm for Monte Carlo simulation of the 2D Ising Model to perform efficiently on a cluster computer using MPI. We use C++ programming language to implement the algorithm. In our algorithm, every…
The paper presents a systematic study and implementation of a reconfigurable combinatorial multi-operand adder for use in Deep Learning systems. The size of carry changes with the number of operands and hence a reliable algorithm to…