English
Related papers

Related papers: Optimized Implementation for Calculation and Fast-…

200 papers

We present a novel approach to accelerate the Goemans-Williamson (GW) randomized rounding procedure for quadratic unconstrained binary optimization (QUBO) problems. Instead of solving the conventional semi-definite programming (SDP)…

Optimization and Control · Mathematics 2025-12-10 Hadi Salloum , Roland Hildebrand , Nhat Trung Nguyen , Vitali Pirau , Amer Al Badr , Mohammad Alkousa , Alexander Gasnikov

We design and develop a work-efficient multithreaded algorithm for sparse matrix-sparse vector multiplication (SpMSpV) where the matrix, the input vector, and the output vector are all sparse. SpMSpV is an important primitive in the…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-10-26 Ariful Azad , Aydin Buluc

We introduce a computationally efficient variant of the model-based ensemble Kalman filter (EnKF). We propose two changes to the original formulation. First, we phrase the setup in terms of precision matrices instead of covariance matrices,…

Methodology · Statistics 2023-03-01 Håkon Gryvill , Håkon Tjelmeland

In this work, we investigate the large-scale mean-field variational inference (MFVI) problem from a mini-batch primal-dual perspective. By reformulating MFVI as a constrained finite-sum problem, we develop a novel primal-dual algorithm…

Machine Learning · Statistics 2026-02-11 Jinhua Lyu , Tianmin Yu , Ying Ma , Naichen Shi

Markov chain Monte Carlo (MCMC) methods are ubiquitous tools for simulation-based inference in many fields but designing and identifying good MCMC samplers is still an open question. This paper introduces a novel MCMC algorithm, namely,…

Despite numerous efforts for optimizing the performance of Sparse Matrix and Vector Multiplication (SpMV) on modern hardware architectures, few works are done to its sparse counterpart, Sparse Matrix and Sparse Vector Multiplication…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-12-18 Min Li , Yulong Ao , Chao Yang

Motivated by an application in high-throughput genomics and metabolomics, we propose a novel, efficient and fully data-driven approach for estimating large block structured sparse covariance matrices in the case where the number of…

Methodology · Statistics 2019-12-09 Marie Perrot-Dockès , Céline Lévy-Leduc , Loïc Rajjou

We present the submatrix method, a highly parallelizable method for the approximate calculation of inverse p-th roots of large sparse symmetric matrices which are required in different scientific applications. We follow the idea of…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-03-06 Michael Lass , Stephan Mohr , Hendrik Wiebeler , Thomas D. Kühne , Christian Plessl

We achieve a quantum speed-up of fully polynomial randomized approximation schemes (FPRAS) for estimating partition functions that combine simulated annealing with the Monte-Carlo Markov Chain method and use non-adaptive cooling schedules.…

Quantum Physics · Physics 2013-06-12 Pawel Wocjan , Chen-Fu Chiang , Anura Abeyesinghe , Daniel Nagaj

This paper describes a fast algorithm for recovering low-rank matrices from their linear measurements contaminated with Poisson noise: the Poisson noise Maximum Likelihood Singular Value thresholding (PMLSV) algorithm. We propose a convex…

Machine Learning · Statistics 2014-12-22 Yang Cao , Yao Xie

In this paper, we present an algorithm for computing a fundamental matrix of formal solutions of completely integrable Pfaffian systems with normal crossings in several variables. This algorithm is a generalization of a method developed for…

Symbolic Computation · Computer Science 2016-10-06 Moulay A. Barkatou , Maximilian Jaroschek , Suzy S. Maddah

We present a multi-block finite-difference solver for massively parallel Direct Numerical Simulations (DNS) of incompressible flows. The algorithm combines the versatility of a multi-block solver with the method of eigenfunctions…

Computational Physics · Physics 2022-09-14 Pedro Costa

Wavefunction correction scheme, which was developed as a variance reduction tool for the pure and fixed-node diffusion Monte Carlo (DMC) computations by Anderson and Freihaut, is applied to the DMC computations of fermions without using the…

Computational Physics · Physics 2010-03-29 Nazim Dugan , Inanc Kanik , Sakir Erkoc

As the performance gains from accelerating quantized matrix multiplication plateau, the softmax operation becomes the critical bottleneck in Transformer inference. This bottleneck stems from two hardware limitations: (1) limited data…

Machine Learning · Computer Science 2026-02-03 Zisheng Ye , Xiaoyu He , Maoyuan Song , Guoliang Qiu , Chao Liao , Chen Wu , Yonggang Sun , Zhichun Li , Xiaoru Xie , Yuanyong Luo , Hu Liu , Pinyan Lu , Heng Liao

In this article we design a novel quasi-regression Monte Carlo algorithm in order to approximate the solution of discrete time backward stochastic differential equations (BSDEs), and we analyze the convergence of the proposed method. The…

Numerical Analysis · Mathematics 2024-08-01 E. Gobet , J. G. López-Salas , C. Vázquez

In this paper, a work-optimal parallelization of Kostelec and Rockmore's well-known fast Fourier transform and its inverse on the three-dimensional rotation group SO(3) is designed, implemented, and tested. To this end, the sequential…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-08-03 Denis-Michael Lux , Christian Wülker , Gregory S. Chirikjian

RISC-V allows for building general-purpose computing platforms with programmable accelerators around a single open-source ISA. However, leveraging heterogeneous SoCs within high-level applications is a tedious task. In this preliminary…

Hardware Architecture · Computer Science 2025-04-08 Cyril Koenig , Enrico Zelioli , Frank K. Gürkaynak , Luca Benini

In a calculation of rotated matrix elements with angular momentum projection, the generalized Wick's theorem may encounter a practical problem of combinatorial complexity when the configurations have more than four quasi-particles (qps).…

Nuclear Theory · Physics 2015-06-22 Long-Jun Wang , Fang-Qi Chen , Takahiro Mizusaki , Makito Oi , Yang Sun

The increasing complexity of AI models requires flexible hardware capable of supporting diverse precision formats, particularly for energy-constrained edge platforms. This work presents PARV-CE, a SIMD-enabled, multi-precision MAC engine…

Hardware Architecture · Computer Science 2025-06-11 Mukul Lokhande , Santosh Kumar Vishvakarma

In this work, we present a new scalable incomplete LU factorization framework called Javelin to be used as a preconditioner for solving sparse linear systems with iterative methods. Javelin allows for improved parallel factorization on…

Mathematical Software · Computer Science 2019-05-06 Joshua Dennis Booth , Gregory Bolet