English
Related papers

Related papers: Extreme Scale FMM-Accelerated Boundary Integral Eq…

200 papers

Due to the high sensitivity of qubits to environmental noise, which leads to decoherence and information loss, active quantum error correction(QEC) is essential. Surface codes represent one of the most promising fault-tolerant QEC schemes,…

Hardware Architecture · Computer Science 2025-07-08 Hao Wang , Erjia Xiao , Wenbo Mu , Songhuan He , Zhongyi Ni , Lingfeng Zhang , Xiaokun Zhan , Yifei Cui , Jinguo Liu , Cheng Wang , Zhongrui Wang , Renjing Xu

We present a solver for the 2D high-frequency Helmholtz equation in heterogeneous acoustic media, with online parallel complexity that scales optimally as $\mathcal{O}(\frac{N}{L})$, where $N$ is the number of volume unknowns, and $L$ is…

Numerical Analysis · Mathematics 2015-08-20 Leonardo Zepeda-Núñez , Laurent Demanet

Modern GPGPUs provide massive arithmetic throughput, yet many scientific kernels remain limited by memory bandwidth. In particular, repeatedly loading precomputed auxiliary data wastes abundant compute resources while stressing the memory…

Performance · Computer Science 2025-11-04 Zijian Cao , Qiao Sun , Tiangong Zhang , Huiyuan Li

Recent hardware acceleration advances have enabled powerful specialized accelerators for finite element computations, spiking neural network inference, and sparse tensor operations. However, existing approaches face fundamental limitations:…

Hardware Architecture · Computer Science 2026-01-09 Chuanzhen Wang , Leo Zhang , Eric Liu

Full-wave 3D electromagnetic simulations of complex planar devices, multilayer interconnects, and chip packages are presented for wide-band frequency-domain analysis using the finite difference integration technique developed in the PETSc…

Computational Engineering, Finance, and Science · Computer Science 2017-05-25 Amir Geranmayeh

The fast proliferation of extreme-edge applications using Deep Learning (DL) based algorithms required dedicated hardware to satisfy extreme-edge applications' latency, throughput, and precision requirements. While inference is achievable…

Hardware Architecture · Computer Science 2022-04-26 Yvan Tortorella , Luca Bertaccini , Davide Rossi , Luca Benini , Francesco Conti

We present a high-order spacetime numerical method for discretizing and solving linear initial-boundary value problems using wavelet-based techniques with user-prescribed error estimates. The spacetime wavelet discretization yields a system…

Numerical Analysis · Mathematics 2025-09-04 Cody D. Cochran , Karel Matous

We examine what is an efficient and scalable nonlinear solver, with low work and memory complexity, for many classes of discretized partial differential equations (PDEs) - matrix-free Full multigrid (FMG) with a Full Approximation Storage…

Numerical Analysis · Mathematics 2023-06-07 Mark F. Adams

We present an accelerated and hardware parallelized integral-equation solver for the problem of acoustic scattering by a two-dimensional surface in three-dimensional space. The approach is based, in part, on the novel Interpolated Factored…

Numerical Analysis · Mathematics 2022-11-01 Edwin Jimenez , Christoph Bauinger , Oscar P. Bruno

In this paper, we discuss the second-order finite element method (FEM) and finite difference method (FDM) for numerically solving elliptic cross-interface problems characterized by vertical and horizontal straight lines, piecewise constant…

Numerical Analysis · Mathematics 2024-11-04 Qiwei Feng

There is a growing interest in custom spatial accelerators for machine learning applications. These accelerators employ a spatial array of processing elements (PEs) interacting via custom buffer hierarchies and networks-on-chip. The…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-06-22 Gordon E. Moon , Hyoukjun Kwon , Geonhwa Jeong , Prasanth Chatarasi , Sivasankaran Rajamanickam , Tushar Krishna

Modern AI accelerators provide high-throughput low-precision matrix engines, but their support for FP32 GEMM is often limited or inefficient. This work presents SGEMM-cube, a precision-recovery FP32 GEMM approximation on Ascend NPUs using…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-07 Weicheng Xue , Baisong Xu , Kai Yang , Yongxiang Liu , Dengdeng Fan , Pengxiang Xu , Yonghong Tian

In this paper we consider a class of robust multilevel precontioners for the Helmholtz equation with high wave number. The key idea in this work is to use the continuous interior penalty finite element methods (CIP-FEM) studied in…

Numerical Analysis · Mathematics 2013-04-26 Huangxin Chen , Haijun Wu , Xuejun Xu

Precise estimation of model inference latency is crucial for time-critical mobile edge applications, enabling devices to calculate latency margins against deadlines and trade them for enhanced model performance or resource savings. However,…

Hardware Architecture · Computer Science 2026-04-20 Jiesong Chen , Jun You , Zhidan Liu , Zhenjiang Li

The boundary element method (BEM) is an efficient numerical method for simulating harmonic wave propagation. It uses boundary integral formulations of the Helmholtz equation at the interfaces of piecewise homogeneous domains. The…

Numerical Analysis · Mathematics 2022-11-01 Elwin van 't Wout , Seyyed R. Haqshenas , Pierre Gélat , Timo Betcke , Nader Saffari

Multiple-input multiple-output orthogonal frequency-division multiplexing (MIMO-OFDM) is a key technology component in the evolution towards cognitive radio (CR) in next-generation communication in which the accuracy of timing and frequency…

Signal Processing · Electrical Eng. & Systems 2022-06-02 Jun Liu , Kai Mei , Xiaochen Zhang , Des McLernon , Dongtang Ma , Jibo Wei , Syed Ali Raza Zaidi

The image quality of the new generation of earthbound Extremely Large Telescopes (ELTs) is heavily influenced by atmospheric turbulences. To compensate these optical distortions a technique called adaptive optics (AO) is used. Many AO…

Numerical Analysis · Mathematics 2020-09-03 Bernadett Stadler , Roberto Biasi , Mauro Manetti , Ronny Ramlau

Top-K SpMV is a key component of similarity-search on sparse embeddings. This sparse workload does not perform well on general-purpose NUMA systems that employ traditional caching strategies. Instead, modern FPGA accelerator cards have a…

Hardware Architecture · Computer Science 2021-03-09 Alberto Parravicini , Luca Giuseppe Cellamare , Marco Siracusa , Marco Domenico Santambrogio

Large language models (LLMs) have demonstrated exceptional proficiency in understanding and generating human language, but efficient inference on resource-constrained embedded devices remains challenging due to large model sizes and…

Hardware Architecture · Computer Science 2025-07-15 Weihong Xu , Haein Choi , Po-kai Hsu , Shimeng Yu , Tajana Rosing

We propose a novel efficient and robust Wavelet-based Edge Multiscale Finite Element Method (WEMsFEM) motivated by \cite{MR3980476,GL18} to solve the singularly perturbed convection-diffusion equations. The main idea is to first establish a…

Numerical Analysis · Mathematics 2024-11-12 Shubin Fu , Eric Chung , Guanglian Li
‹ Prev 1 4 5 6 7 8 10 Next ›