Related papers: Parallel Implementation of the PHOENIX Generalized…
Radiative transfer (RT) modelling is a necessary tool in the interpretation of observations of the thermal emission of interstellar dust. It is also often part of multi-physics modelling. In this context, the efficiency of radiative…
Optical networks with parallel processing capabilities are significant in advancing high-speed data computing and large-scale data processing by providing ultra-width computational bandwidth. In this paper, we present a photonic integrated…
Computational chemistry allows researchers to experiment in sillico: by running a computer simulations of a biological or chemical processes of interest. Molecular dynamics with molecular mechanics model of interactions simulates N-body…
Tensor processing is the cornerstone of modern technological advancements, powering critical applications in data analytics and artificial intelligence. While optical computing offers exceptional advantages in bandwidth, parallelism, and…
Video diffusion models (VDMs) perform attention computation over the 3D spatio-temporal domain. Compared to large language models (LLMs) processing 1D sequences, their memory consumption scales cubically, necessitating parallel serving…
Metasurfaces are nano-structured devices composed of arrays of subwavelength scatterers (or meta-atoms) that manipulate the wavefront, polarization, or intensity of light. Like other diffractive optical devices, metasurfaces suffer from…
Context. Stellar spectral synthesis is essential for various applications, ranging from determining stellar parameters to comprehensive stellar variability calculations. New observational resources as well as advanced stellar atmosphere…
Accurate and fast modeling of electric fields in layered structures have a great scientific and practical value. Prevalent method for that is transfer-matrix method. However, transfer matrix method is limited to infinite plane wave…
Programmable photonic circuits are versatile platforms that route light through multiple interference paths using reconfigurable optoelectronic elements to perform complex discrete linear operations. These circuits offer the potential for…
The development of deep neural networks is witnessing fast growth in network size, which requires novel hardware computing platforms with large bandwidth and low energy consumption. Optical computing has been a potential candidate for…
Kernel matrix-vector product is ubiquitous in many science and engineering applications. However, a naive method requires $O(N^2)$ operations, which becomes prohibitive for large-scale problems. We introduce a parallel method that provably…
In the near future, massively parallel computing systems will be necessary to solve computation intensive applications. The key bottleneck in massively parallel implementation of numerical algorithms is the synchronization of data across…
We describe a program for the parallel implementation of multiple runs of XSTAR, a photoionization code that is used to predict the physical properties of an ionized gas from its emission and/or absorption lines. The parallelization…
We propose a parallel protocol for implementing distributed nonlocal quantum gates between spatially separated stationary qubits encoded in dual-species quantum emitters (i.e., color-center and superconducting qubits). By utilizing…
We present a parallel FFT algorithm for SIMD systems following the `Transpose Algorithm' approach. The method is based on the assignment of the data field onto a 1-dimensional ring of systolic cells. The systolic array can be universally…
This report discusses the implementation of two parallel algorithms on a distributed memory system for studying vortex dynamics in type-II superconductors. These algorithms are the same as that implemented for classical molecular dynamics…
Optical computing represents a groundbreaking technology that leverages the unique properties of photons, with innate parallelism standing as its most compelling advantage. Parallel optical computing like cascaded Mach-Zehnder…
Communication is a key bottleneck for distributed graph neural network (GNN) training. This paper proposes GNNPipe, a new approach that scales the distributed full-graph deep GNN training. Being the first to use layer-level model…
In this article we consider the inversion problem for polynomially computable discrete functions. These functions describe behavior of many discrete systems and are used in model checking, hardware verification, cryptanalysis, computer…
High-speed signal processing is essential for maximizing data throughput in emerging communication applications, like multiple-input multiple-output (MIMO) systems and radio-frequency (RF) interference cancellation. However, as these…