English
Related papers

Related papers: Advanced Techniques for High-Performance Fock Matr…

200 papers

Lattice spin models are useful for studying critical phenomena and allow the extraction of equilibrium and dynamical properties. Simulations of such systems are usually based on Monte Carlo (MC) techniques, and the main difficulty is often…

Computational Physics · Physics 2012-09-13 Tal Levy , Guy Cohen , Eran Rabani

Robust trajectory optimization enables autonomous systems to operate safely under uncertainty by computing control policies that satisfy the constraints for all bounded disturbances. However, these problems often lead to large Second Order…

Robotics · Computer Science 2026-05-19 Jiawei Wang , Arshiya Taj Abdul , Evangelos A. Theodorou

Among the algorithms that are likely to play a major role in future exascale computing, the fast multipole method (FMM) appears as a rising star. Our previous recent work showed scaling of an FMM on GPU clusters, with problem sizes in the…

Numerical Analysis · Computer Science 2012-10-30 Rio Yokota , Lorena Barba

This paper introduces a framework for solving alternating current optimal power flow (ACOPF) problems using graphics processing units (GPUs). While GPUs have demonstrated remarkable performance in various computing domains, their…

Optimization and Control · Mathematics 2026-05-11 Sungho Shin , François Pacaud , Mihai Anitescu

Multiple matching algorithms are used to locate the occurrences of patterns from a finite pattern set in a large input string. Aho-Corasick and Wu-Manber, two of the most well known algorithms for multiple matching require an increased…

Distributed, Parallel, and Cluster Computing · Computer Science 2014-07-11 Charalampos S. Kouzinopoulos , John-Alexander M. Assael , Themistoklis K. Pyrgiotis , Konstantinos G. Margaritis

We present Occamy, a 432-core RISC-V dual-chiplet 2.5D system for efficient sparse linear algebra and stencil computations on FP64 and narrow (32-, 16-, 8-bit) SIMD FP data. Occamy features 48 clusters of RISC-V cores with custom…

We describe an interface and an implementation for performing Kronecker product actions on NVIDIA GPUs for multiple small 2-D matrices and 3-D arrays processed in parallel as a batch. This method is suited to cases where the Kronecker…

Mathematical Software · Computer Science 2013-04-29 Chetan Jhurani

Spectral clustering is one of the most popular graph clustering algorithms, which achieves the best performance for many scientific and engineering applications. However, existing implementations in commonly used software platforms such as…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-02-14 Yu Jin , Joseph F. JaJa

Low-dose Proton Computed Tomography (pCT) is an evolving imaging modality that is used in proton therapy planning which addresses the range uncertainty problem. The goal of pCT is generating a 3D map of Relative Stopping Power (RSP)…

Matrix multiplication is a fundamental operation in both training of neural networks and inference. To accelerate matrix multiplication, Graphical Processing Units (GPUs) provide it implemented in hardware. Due to the increased throughput…

Mathematical Software · Computer Science 2026-04-07 Faizan A. Khattak , Mantas Mikaitis

Branch-and-Bound (B&B) algorithms are time intensive tree-based exploration methods for solving to optimality combinatorial optimization problems. In this paper, we investigate the use of GPU computing as a major complementary way to speed…

Distributed, Parallel, and Cluster Computing · Computer Science 2012-08-21 Melab Nouredine , Imen Chakroun , Mezmaz Mohand , Daniel Tuyttens

Massive multi-threading in GPU imposes tremendous pressure on memory subsystems. Due to rapid growth in thread-level parallelism of GPU and slowly improved peak memory bandwidth, the memory becomes a bottleneck of GPU's performance and…

Hardware Architecture · Computer Science 2019-06-17 Bing Li , Mengjie Mao , Xiaoxiao Liu , Tao Liu , Zihao Liu , Wujie Wen , Yiran Chen , Hai , Li

Novel methods are presented in this initial study for the fusion of GPU kernels in the artificial compressibility method (ACM), using tensor product elements with constant Jacobians and flux reconstruction. This is made possible through the…

Mathematical Software · Computer Science 2022-01-05 Will Trojak , Rob Watson , Freddie Witherden

This paper presents a methodology for simultaneous heterogeneous computing, named ENEAC, where a quad core ARM Cortex-A53 CPU works in tandem with a preprogrammed on-board FPGA accelerator. A heterogeneous scheduler distributes the tasks…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-11-16 Kris Nikov , Mohammad Hosseinabady , Rafael Asenjo , Andrés Rodríguezz , Angeles Navarro , Jose Nunez-Yanez

A linear-scaling algorithm is presented for computing the Hartree-Fock (HF) exchange matrix using concentric atomic density fitting. The algorithm utilizes the stronger distance dependence of the three-center electron repulsion integrals…

Chemical Physics · Physics 2014-10-21 David S. Hollman , Henry F. Schaefer , Edward F. Valeev

We investigate the performance of Opticks, a NVIDIA OptiX API 7.5 GPU-accelerated photon propagation tool compared with a single-threaded Geant4 simulation. We compare the simulations using an improved model of the NEXT-CRAB-0 gaseous time…

Instrumentation and Detectors · Physics 2025-11-20 NEXT Collaboration , I. Parmaksiz , K. Mistry , E. Church , C. Adams , J. Asaadi , J. Baeza-Rubio , K. Bailey , N. Byrnes , B. J. P. Jones , I. A. Moya , K. E. Navarro , D. R. Nygren , P. Oyedele , L. Rogers , F. Samaniego , K. Stogsdill , H. Almazán , V. Álvarez , B. Aparicio , A. I. Aranburu , L. Arazi , I. J. Arnquist , F. Auria-Luna , S. Ayet , C. D. R. Azevedo , F. Ballester , M. del Barrio-Torregrosa , A. Bayo , J. M. Benlloch-Rodríguez , F. I. G. M. Borges , A. Brodolin , S. Cárcel , A. Castillo , L. Cid , C. A. N. Conde , T. Contreras , F. P. Cossío , R. Coupe , E. Dey , G. Díaz , C. Echevarria , M. Elorza , J. Escada , R. Esteve , R. Felkai , L. M. P. Fernandes , P. Ferrario , A. L. Ferreira , F. W. Foss , Z. Freixa , J. García-Barrena , J. J. Gómez-Cadenas , J. W. R. Grocott , R. Guenette , J. Hauptman , C. A. O. Henriques , J. A. Hernando Morata , P. Herrero-Gómez , V. Herrero , C. Hervés Carrete , Y. Ifergan , F. Kellerer , L. Larizgoitia , A. Larumbe , P. Lebrun , F. Lopez , N. López-March , R. Madigan , R. D. P. Mano , A. P. Marques , J. Martín-Albo , G. Martínez-Lema , M. Martínez-Vara , R. L. Miller , J. Molina-Canteras , F. Monrabal , C. M. B. Monteiro , F. J. Mora , P. Novella , A. Nuñez , E. Oblak , J. Palacio , B. Palmeiro , A. Para , A. Pazos , J. Pelegrin , M. Pérez Maneiro , M. Querol , J. Renner , I. Rivilla , C. Rogero , B. Romeo , C. Romo-Luque , V. San Nacienciano , F. P. Santos , J. M. F. dos Santos , M. Seemann , I. Shomroni , P. A. O. C. Silva , A. Simón , S. R. Soleti , M. Sorel , J. Soto-Oton , J. M. R. Teixeira , S. Teruel-Pardo , J. F. Toledo , C. Tonnelé , S. Torelli , J. Torrent , A. Trettin , A. Usón , P. R. G. Valle , J. F. C. A. Veloso , J. Waiton , A. Yubero-Navarro

Optical flow estimation is crucial for autonomous navigation and localization of unmanned aerial vehicles (UAV). On micro and nano UAVs, real-time calculation of the optical flow is run on low power and resource-constrained microcontroller…

Computer Vision and Pattern Recognition · Computer Science 2023-05-23 Jonas Kühne , Michele Magno , Luca Benini

In this paper, we propose the first optimum process scheduling algorithm for an increasingly prevalent type of heterogeneous multicore (HEMC) system that combines high-performance big cores and energy-efficient small cores with the same…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-09-13 Chien-Hao Chen , Ren-Song Tsay

As in various fields like scientific research and industrial application, the computation time optimization is becoming a task that is of increasing importance because of its highly parallel architecture. The graphics processing unit is…

Performance · Computer Science 2017-10-18 Huichao Hong , Lixin Zheng , Shuwan Pan

In this paper, we describe the architecture and performance of the GraCCA system, a Graphic-Card Cluster for Astrophysics simulations. It consists of 16 nodes, with each node equipped with 2 modern graphic cards, the NVIDIA GeForce 8800…

Astrophysics · Physics 2008-11-26 Hsi-Yu Schive , Chia-Hung Chien , Shing-Kwong Wong , Yu-Chih Tsai , Tzihong Chiueh