English
Related papers

Related papers: A SVD accelerated kernel-independent fast multipol…

200 papers

Transmission matrix (TM) linearly maps the incident and transmitted complex fields, and has been used widely due to its ability to characterize scattering media. It is computationally demanding to reconstruct the TM from intensity images…

Optics · Physics 2024-01-05 Jingshan Zhong , Zhong Wen , Quanzhi Li , Qilin Deng , Qing Yang

Large multimodal models (LMMs) suffer significant computational challenges due to the high cost of Large Language Models (LLMs) and the quadratic complexity of processing long vision token sequences. In this paper, we explore the spatial…

Computer Vision and Pattern Recognition · Computer Science 2025-06-10 Hao Tang , Chengchao Shen

We present VitaLLM, a mixed precision accelerator that enables ternary weight large language models to run efficiently on edge devices. The design combines two compute cores, a multiplier free TINT core for ternary-INT projections and a…

Hardware Architecture · Computer Science 2026-05-04 Zi-Wei Lin , Tian-Sheuan Chang

Kernel adaptive filtering (KAF) integrates traditional linear algorithms with kernel methods to generate nonlinear solutions in the input space. The standard approach relies on the representer theorem and the kernel trick to perform…

Signal Processing · Electrical Eng. & Systems 2025-01-16 Kan Li , Jose C. Principe

With applications ranging from metabolomics to histopathology, quantitative phase microscopy (QPM) is a powerful label-free imaging modality. Despite significant advances in fast multiplexed imaging sensors and deep-learning-based inverse…

We present an algorithm to parallelize the inverse fast multipole method (IFMM), which is an approximate direct solver for dense linear systems. The parallel scheme is based on a greedy coloring algorithm, where two nodes in the hierarchy…

Computational Physics · Physics 2020-02-19 Toru Takahashi , Chao Chen , Eric Darve

Characteristic mode (CM) analysis poses challenges in computational electromagnetics (CEM) as it calls for efficient solutions of dense generalized eigenvalue problems (GEP). Multilevel fast multipole algorithm (MLFMA) can greatly reduce…

Numerical Analysis · Mathematics 2014-12-10 Qi I. Dai , Jun Wei Wu , Ling Ling Meng , Weng Cho Chew , Wei E. I. Sha

In this study, a fast multipole method (FMM) is used to decrease the computational time of a fully-coupled poroelastic hydraulic fracture model with a controllable effect on its accuracy. The hydraulic fracture model is based on the…

Numerical Analysis · Computer Science 2019-10-23 Ali Rezaei , Fahd Siddiqui , Giorgio Bornia , Mohamed Y. Soliman

An important linear algebra routine, GEneral Matrix Multiplication (GEMM), is a fundamental operator in deep learning. Compilers need to translate these routines into low-level code optimized for specific hardware. Compiler-level…

Machine Learning · Computer Science 2019-09-25 Huaqing Zhang , Xiaolin Cheng , Hui Zang , Dae Hoon Park

With the assistance of singular value decomposition (SVD), a multi-beam directional modulation (DM) scheme based on symmetrical multi-carrier frequency diverse array (FDA) is proposed. The proposed DM scheme is capable of achieving…

Signal Processing · Electrical Eng. & Systems 2019-07-30 Qian Cheng , Vincent Fusco , Jiang Zhu , Shilian Wang , Chao Gu

This paper proposes a novel multimodal deep learning framework integrating bidirectional LSTM, multi-head attention mechanism, and variational mode decomposition (BiLSTM-AM-VMD) for early liver cancer diagnosis. Using heterogeneous data…

Machine Learning · Computer Science 2025-09-03 Cheng Cheng , Zeping Chen , Xavier Wang

Discrete cosine transform (DCT) and other Fourier-related transforms have broad applications in scientific computing. However, off-the-shelf high-performance multi-dimensional DCT (MD DCT) libraries are not readily available in parallel…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-10-05 Zixuan Jiang , Jiaqi Gu , David Z. Pan

In this paper, we report the results obtained from the acceleration of multi-binary64-type multiple precision matrix multiplication with AVX2. We target double-double (DD), triple-double (TD), and quad-double (QD) precision arithmetic…

Numerical Analysis · Mathematics 2021-09-14 Tomonori Kouya

Nonnegative matrix factorization (NMF) is a powerful technique for dimension reduction, extracting latent factors and learning part-based representation. For large datasets, NMF performance depends on some major issues: fast algorithms,…

Optimization and Control · Mathematics 2015-07-01 Duy-Khuong Nguyen , Tu-Bao Ho

Spectral embedding based on the Singular Value Decomposition (SVD) is a widely used "preprocessing" step in many learning tasks, typically leading to dimensionality reduction by projecting onto a number of dominant singular vectors and…

Machine Learning · Statistics 2015-09-29 Dinesh Ramasamy , Upamanyu Madhow

Among optimal hierarchical algorithms for the computational solution of elliptic problems, the Fast Multipole Method (FMM) stands out for its adaptability to emerging architectures, having high arithmetic intensity, tunable accuracy, and…

Numerical Analysis · Computer Science 2016-01-20 Huda Ibeid , Rio Yokota , Jennifer Pestana , David Keyes

Symmetric positive semi-definite (SPSD) matrix approximation methods have been extensively used to speed up large-scale eigenvalue computation and kernel learning methods. The standard sketch based method, which we call the prototype model,…

Machine Learning · Computer Science 2016-12-13 Shusen Wang , Zhihua Zhang , Tong Zhang

The deployment of large language models (LLMs) is often constrained by memory bandwidth, where the primary bottleneck is the cost of transferring model parameters from the GPU's global memory to its registers. When coupled with custom…

Machine Learning · Computer Science 2025-01-20 Han Guo , William Brandon , Radostin Cholakov , Jonathan Ragan-Kelley , Eric P. Xing , Yoon Kim

Recently, a new framework to compute the photoionization rate in streamer discharges accurately and efficiently using the integral form and the fast multipole method (FMM) was presented. This paper further improves the efficiency of this…

Plasma Physics · Physics 2021-12-21 Bo Lin , Chijie Zhuang

Mixed-precision quantization is a promising approach for compressing large language models under tight memory budgets. However, existing mixed-precision methods typically suffer from one of two limitations: they either rely on expensive…

Machine Learning · Computer Science 2026-02-03 Xin Nie , Haicheng Zhang , Liang Dong , Beining Feng , Jinhong Weng , Guiling Sun
‹ Prev 1 4 5 6 7 8 10 Next ›