Related papers: Accelerating Dedispersion using Many-Core Architec…
There has been tremendous progress in the physical realization of quantum computing hardware in recent times, bringing us closer than ever before to realizing the promise of quantum computing. However, noise continues to pose a crucial…
The dispersion of extragalactic fast radio bursts (FRBs) can serve as a powerful probe of the diffuse plasma between and surrounding galaxies, which contains most of the Universe's baryons. By cross-correlating the dispersion of background…
The expansion of telecommunications incurs increasingly severe crosstalk and interference, and a physical layer cognitive method, called blind source separation (BSS), can effectively address these issues. BSS requires minimal prior…
Multi-core neuromorphic processors are becoming increasingly significant due to their energy-efficient local computing and scalable modular architecture, particularly for event-based processing applications. However, minimizing the cost of…
We present a high-performance, graphics processing unit (GPU)-based framework for the efficient analysis and visualization of (nearly) terabyte (TB)-sized 3-dimensional images. Using a cluster of 96 GPUs, we demonstrate for a 0.5 TB image:…
We present the methodology of a photon-conserving, spatially-adaptive, ray-tracing radiative transfer algorithm, designed to run on multiple parallel Graphic Processing Units (GPUs). Each GPU has thousands computing cores, making them…
Modern supercomputers are increasingly requiring the presence of accelerators and co-processors. However, it has not been easy to achieve good performance on such heterogeneous clusters. The key challenge has been to ensure good load…
Digital deblurring of images is an important problem that arises in multifrequency observations of the Cosmic Microwave Background (CMB) where, because of the width of the point spread functions (PSF), maps at different frequencies suffer a…
Recently, diffusion models (DMs) have made significant strides in high-quality image generation. However, the multi-step denoising process often results in considerable computational overhead, impeding deployment on resource-constrained…
Dispersive Fourier transformation is a powerful technique in which the spectrum of an optical pulse is mapped into a time-domain waveform using chromatic dispersion. It replaces a diffraction grating and detector array with a dispersive…
In this paper we implemented the algorithm we developed in [1] called 3DPIFCM in a parallel environment by using CUDA on a GPU. In our previous work we introduced 3DPIFCM which performs segmentation of images in noisy conditions and uses…
Physics-inspired and quantum compute based methods for processing in the physical layer of next-generation cellular radio access networks have demonstrated theoretical advances in spectral efficiency in recent years, but have stopped short…
Fast Radio Bursts are millisecond bursts of radio radiation at frequencies of about 1 GHz, recently discovered in pulsar surveys. They have not yet been definitively identified with any other astronomical object or phenomenon. The bursts…
High level programming languages and GPU accelerators are powerful enablers for a wide range of applications. Achieving scalable vertical (within a compute node), horizontal (across compute nodes), and temporal (over different generations…
Discrete optimization is a central problem in artificial intelligence. The optimization of the aggregated cost of a network of cost functions arises in a variety of problems including (W)CSP, DCOP, as well as optimization in stochastic…
In modern wireless networks, interference is no longer negligible since each cell becomes smaller to support high throughput. The reduced size of each cell forces to install many cells, and consequently causes to increase inter-cell…
Emerging analog computing substrates, such as oscillator-based Ising machines, offer rapid convergence times for combinatorial optimization but often suffer from limited scalability due to physical implementation constraints. To tackle…
This article introduces a highly parallel algorithm for molecular dynamics simulations with short-range forces on single node multi- and many-core systems. The algorithm is designed to achieve high parallel speedups for strongly…
Among error-correcting codes, polar codes are the first to provably achieve channel capacity with an explicit construction. In this work, we present software implementations of a polar decoder that leverage the capabilities of modern…
In the next decade, the demands for computing in large scientific experiments are expected to grow tremendously. During the same time period, CPU performance increases will be limited. At the CERN Large Hadron Collider (LHC), these two…