English
Related papers

Related papers: Accelerating Fock build via hybrid analytical-nume…

200 papers

Fast Fourier Transform (FFT) is an essential tool in scientific and engineering computation. The increasing demand for mixed-precision FFT has made it possible to utilize half-precision floating-point (FP16) arithmetic for faster speed and…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-04-26 Binrui Li , Shenggan Cheng , James Lin

We set new speed records for multiplying long polynomials over finite fields of characteristic two. Our multiplication algorithm is based on an additive FFT (Fast Fourier Transform) by Lin, Chung, and Huang in 2014 comparing to previously…

Symbolic Computation · Computer Science 2018-01-08 Ming-Shing Chen , Chen-Mou Cheng , Po-Chun Kuo , Wen-Ding Li , Bo-Yin Yang

There has been considerable research into improving Fast Fourier Transform (FFT) performance through parallelization and optimization for specialized hardware. However, even with those advancements, processing of very large files, over 1TB…

Distributed, Parallel, and Cluster Computing · Computer Science 2014-07-28 Rostislav Tsiomenko , Bradley S. Rees

Analog computing has been recognized as a promising low-power alternative to digital counterparts for neural network acceleration. However, conventional analog computing is mainly in a mixed-signal manner. Tedious analog/digital (A/D)…

Emerging Technologies · Computer Science 2022-08-18 Hanqing Zhu , Keren Zhu , Jiaqi Gu , Harrison Jin , Ray Chen , Jean Anne Incorvia , David Z. Pan

Fine-pitch hybridisation processes are essential for next-generation pixel detectors and high-density microelectronic assemblies. Conventional bump-bonding techniques, although reliable, remain costly and difficult to implement for…

Analog Compute-in-Memory (CiM) accelerators are increasingly recognized for their efficiency in accelerating Deep Neural Networks (DNN). However, their dependence on Analog-to-Digital Converters (ADCs) for accumulating partial sums from…

Hardware Architecture · Computer Science 2024-03-21 Shubham Negi , Utkarsh Saxena , Deepika Sharma , Kaushik Roy

Monte Carlo Application Toolkit (MCATK) commonly uses surface tracking on a structured mesh to compute scalar fluxes. In this mode, higher fidelity requires more mesh cells and isotopes and thus more computational overhead -- since every…

Computational Physics · Physics 2023-06-14 J. P. Morgan , Travis J. Trahan , Timothy P. Burke , Colin J. Josey , Kyle E. Niemeyer

This work presents a multi-layered methodology for efficiently accelerating multimodal foundation models (MFMs). It combines hardware and software co-design of transformer blocks with an optimization pipeline that reduces computational and…

On modern architectures, the performance of 32-bit operations is often at least twice as fast as the performance of 64-bit operations. By using a combination of 32-bit and 64-bit floating point arithmetic, the performance of many dense and…

Mathematical Software · Computer Science 2015-05-13 Marc Baboulin , Alfredo Buttari , Jack Dongarra , Jakub Kurzak , Julie Langou , Julien Langou , Piotr Luszczek , Stanimire Tomov

At least 25 kinds of detector-like devices need to be integrated in Phase I of the High Energy Photon Source (HEPS), and the work needs to be carefully planned to maximise productivity with highly limited human resources. After a systematic…

Instrumentation and Detectors · Physics 2024-11-06 Qun Zhang , Peng-Cheng Li , Ling-Zhu Bian , Chun Li , Zong-Yang Yue , Cheng-Long Zhang , Zhuo-Feng Zhao , Yi Zhang , Gang Li , Ai-Yu Zhou , Yu Liu

VoxCap, a fast Fourier transform (FFT)-accelerated and Tucker-enhanced integral equation simulator for capacitance extraction of voxelized structures, is proposed. The VoxCap solves the surface integral equations (SIEs) for conductor and…

Computational Engineering, Finance, and Science · Computer Science 2021-02-24 Mingyu Wang , Cheng Qian , Jacob K. White , Abdulkadir C. Yucel

Anchor-based methods are a pivotal approach in handling clustering of large-scale data. However, these methods typically entail two distinct stages: selecting anchor points and constructing an anchor graph. This bifurcation, along with the…

Machine Learning · Computer Science 2024-10-31 Shikun Mei , Fangfang Li , Quanxue Gao , Ming Yang

This paper presents an accuracy-enhanced Hybrid Temporal Computing (E-HTC) framework for ultra-low-power hardware accelerators with deterministic additions. Inspired by the recently proposed HTC architecture, which leverages pulse-rate and…

Hardware Architecture · Computer Science 2025-09-30 Sachin Sachdeva , Jincong Lu , Wantong Li , Sheldon X. -D. Tan

We introduce a practical hybrid approach that combines orbital-free density functional theory (DFT) with Kohn-Sham DFT for speeding up first-principles molecular dynamics simulations. Equilibrated ionic configurations are generated using…

We present a novel low latency CMOS hardware accelerator for fully connected (FC) layers in deep neural networks (DNNs). The FC accelerator, FC-ACCL, is based on 128 8x8 or 16x16 processing elements (PEs) for matrix-vector multiplication,…

Hardware Architecture · Computer Science 2020-11-26 Nick Iliev , Amit Ranjan Trivedi

Hybrid density functional theory (DFT) remains intractable for large periodic systems due to the demanding computational cost of exact exchange. We apply the tensor hypercontraction (THC) (or interpolative separable density fitting)…

Computational Physics · Physics 2023-10-13 Adam Rettig , Joonho Lee , Martin Head-Gordon

We present a computationally efficient approach to perform systematically convergent real-space all-electron Kohn-Sham DFT calculations for solids using an enriched finite element (FE) basis. The enriched FE basis is constructed by…

Computational Physics · Physics 2021-08-18 Nelson D. Rufus , Bikash Kanungo , Vikram Gavini

The edge processing of deep neural networks (DNNs) is becoming increasingly important due to its ability to extract valuable information directly at the data source to minimize latency and energy consumption. Frequency-domain model…

Hardware Architecture · Computer Science 2023-09-06 Nastaran Darabi , Maeesha Binte Hashem , Hongyi Pan , Ahmet Cetin , Wilfred Gomes , Amit Ranjan Trivedi

Recent hardware-aware matrix-free algorithms for higher-order finite-element (FE) discretized matrix-vector multiplications reduce floating point operations and data access costs compared to traditional sparse matrix approaches. This work…

Computational Physics · Physics 2024-12-31 Gourab Panigrahi , Nikhil Kodali , Debashis Panda , Phani Motamarri

We investigate the use of optimized correlation consistent gaussian basis sets for the study of insulating solids with auxiliary-field quantum Monte Carlo (AFQMC). The exponents of the basis set are optimized through the minimization of the…

Chemical Physics · Physics 2021-02-03 Miguel A. Morales , Fionn D. Malone
‹ Prev 1 3 4 5 6 7 10 Next ›