English
Related papers

Related papers: Fast distributed phononic band-structure calculati…

200 papers

We develop fast approximation algorithms for the minimum-cost version of the Bounded-Degree MST problem (BD-MST) and its generalization the Crossing Spanning Tree problem (Crossing-ST). We solve the underlying LP to within a $(1+\epsilon)$…

Data Structures and Algorithms · Computer Science 2021-05-19 Chandra Chekuri , Kent Quanrud , Manuel R. Torres

The Kernel Polynomial Method (KPM) is one of the fast diagonalization methods used for simulations of quantum systems in research fields of condensed matter physics and chemistry. The algorithm has a difficulty to be parallelized on a…

Computational Physics · Physics 2011-05-30 Shixun Zhang , Shinichi Yamagiwa , Masahiko Okumura , Seiji Yunoki

Space group theory is pivotal in the design of nanophotonics devices, enabling the characterization of periodic optical structures such as photonic crystals. The aim of this study is to extend the application of nonsymmorphic space groups…

Optics · Physics 2025-06-02 Lida Liu , Jingwei Wang , Yuhao Jing , Songzi Lin , Zhongfei Xiong , Yuntian Chen

This paper describes the main features of a pioneering unsteady solver for simulating ideal two-fluid plasmas on unstructured grids, taking profit of GPGPU (General-purpose computing on graphics processing units). The code, which has been…

Accelerating the deep learning inference is very important for real-time applications. In this paper, we propose a novel method to fuse the layers of convolutional neural networks (CNNs) on Graphics Processing Units (GPUs), which applies…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-07-30 Xueying Wang , Guangli Li , Xiao Dong , Jiansong Li , Lei Liu , Xiaobing Feng

The Graphics Processing Unit (GPU) has become an integral part of astronomical instrumentation, enabling high-performance online data reduction and accelerated online signal processing. In this paper, we describe a wide-band reconfigurable…

Instrumentation and Methods for Astrophysics · Physics 2015-06-23 Jayanth Chennamangalam , Simon Scott , Glenn Jones , Hong Chen , John Ford , Amanda Kepley , D. R. Lorimer , Jun Nie , Richard Prestage , D. Anish Roshi , Mark Wagner , Dan Werthimer

A novel and scalable geometric multi-level algorithm is presented for the numerical solution of elliptic partial differential equations, specially designed to run with high occupancy of streaming processors inside Graphics Processing…

Mathematical Software · Computer Science 2017-03-22 J. T. Becerra-Sagredo , F. Mandujano , C. Malaga

GPUs have significantly accelerated first-order methods for large-scale optimization, especially in continuous optimization. However, this success has not transferred cleanly to problems with discrete variables, combinatorial structure, and…

Machine Learning · Computer Science 2026-05-22 Jiachang Liu , Andrea Lodi

Designing photonic circuits that prepare graph states with high fidelity and success probability is a central challenge in linear optical quantum computing. Existing approaches rely on hand-crafted designs or fusion-based assemblies. In the…

We present a general method for accelerating by more than an order of magnitude the convolution of pixelated function on the sphere with a radially-symmetric kernel. Our method splits the kernel into a compact real-space, and a compact…

Instrumentation and Methods for Astrophysics · Physics 2012-11-16 P. M. Sutter , Benjamin D. Wandelt , Franz Elsner

Graphics Processing Units (GPUs) are high performance co-processors originally intended to improve the use and quality of computer graphics applications. Once, researchers and practitioners noticed the potential of using GPU for general…

Numerical Analysis · Computer Science 2016-07-12 K. Parand , Saeed Zafarvahedian , Sayyed A. Hossayni

This paper presents a novel boundary-optimized fast Fourier extension algorithm for efficient approximation of non-periodic functions. The proposed methodology constructs periodic extensions through strategic utilization of boundary…

Numerical Analysis · Mathematics 2025-08-27 Z. Y. Zhao , Y. F Wang , A. G. Yagola

Deformable image registration and regression are important tasks in medical image analysis. However, they are computationally expensive, especially when analyzing large-scale datasets that contain thousands of images. Hence, cluster…

Computer Vision and Pattern Recognition · Computer Science 2017-11-17 Zhipeng Ding , Greg Fleishman , Xiao Yang , Paul Thompson , Roland Kwitt , Marc Niethammer

Recently, practical analog in-memory computing has been realized using unmodified commercial DRAM modules. The underlying Processing-Using-DRAM (PUD) techniques enable high-throughput bitwise operations directly within DRAM arrays. However,…

Hardware Architecture · Computer Science 2025-05-09 Tatsuya Kubo , Daichi Tokuda , Lei Qu , Ting Cao , Shinya Takamaeda-Yamazaki

We present a GPU-accelerated cosmological simulation code, PhotoNs-GPU, based on algorithm of Particle Mesh Fast Multipole Method (PM-FMM), and focus on the GPU utilization and optimization. A proper interpolated method for truncated…

Instrumentation and Methods for Astrophysics · Physics 2021-12-28 Qiao Wang , Chen Meng

This work introduces a kernel-independent, multilevel, adaptive algorithm for efficiently evaluating a discrete convolution kernel with a given source distribution. The method is based on linear algebraic tools such as low rank…

Numerical Analysis · Mathematics 2025-07-11 Anna Yesypenko , Chao Chen , Per-Gunnar Martinsson

The exponential growth of artificial intelligence has fueled the development of high-bandwidth photonic interconnect fabrics as a critical component of modern AI supercomputers. As the demand for ever-increasing AI compute and connectivity…

Optics · Physics 2025-03-03 Jesse Lu , David Qu , Jim Qu , Ryan Fong , Geun Ho Ahn , Jelena Vuckovic

In this paper, we present the details of our multi-node GPU-FFT library, as well its scaling on Selene HPC system. Our library employs slab decomposition for data division and MPI for communication among GPUs. We performed GPU-FFT on…

In this work we propose an accelerated stochastic learning system for very large-scale applications. Acceleration is achieved by mapping the training algorithm onto massively parallel processors: we demonstrate a parallel, asynchronous GPU…

Machine Learning · Computer Science 2017-02-24 Thomas Parnell , Celestine Dünner , Kubilay Atasu , Manolis Sifalakis , Haris Pozidis

Exactly computing the full output distribution of linear optical circuits remains a challenge, as existing methods are either time-efficient but memory-intensive or memory-efficient but slow. Moreover, any realistic simulation must account…

Quantum Physics · Physics 2025-03-10 Timothée Goubault de Brugière , Nicolas Heurtel
‹ Prev 1 3 4 5 6 7 10 Next ›