English
Related papers

Related papers: GPU-accelerated finite-temperature Lanczos method …

200 papers

Weighted finite automata and transducers (including hidden Markov models and conditional random fields) are widely used in natural language processing (NLP) to perform tasks such as morphological analysis, part-of-speech tagging, chunking,…

Computation and Language · Computer Science 2017-01-18 Arturo Argueta , David Chiang

We present a novel finite element integration method for low order elements on GPUs. We achieve more than 100GF for element integration on first order discretizations of both the Laplacian and Elasticity operators.

Mathematical Software · Computer Science 2014-09-29 Matthew G. Knepley , Andy R. Terrel

A modern graphics processing unit (GPU) is able to perform massively parallel scientific computations at low cost. We extend our implementation of the checkerboard algorithm for the two dimensional Ising model [T. Preis et al., J. Comp.…

Computational Physics · Physics 2010-07-22 Benjamin Block , Peter Virnau , Tobias Preis

Computing on graphics processors is maybe one of the most important developments in computational science to happen in decades. Not since the arrival of the Beowulf cluster, which combined open source software with commodity hardware to…

Mathematical Software · Computer Science 2011-09-21 Felipe A. Cruz , Simon K. Layton , Lorena A. Barba

RAPID-LLM is a unified performance modeling framework for large language model (LLM) training and inference on GPU clusters. It couples a DeepFlow-based frontend that generates hardware-aware, operator-level Chakra execution traces from an…

We review a recent approach for the simulation of many-body interacting systems based on an efficient generalization of the Lanczos method for Quantum Monte Carlo simulations. This technique allows to perform systematic corrections to a…

Strongly Correlated Electrons · Physics 2007-05-23 Sandro Sorella

The rise of Generative AI introduces a new class of HPC workloads that integrates lightweight LLMs with traditional high-throughput applications to accelerate scientific discovery. The current design of HPC clusters is inadequate to support…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-17 Thanh Son Phung , Douglas Thain

In this thesis we develop techniques to efficiently solve numerical Partial Differential Equations (PDEs) using Graphical Processing Units (GPUs). Focus is put on both performance and re--usability of the methods developed, to this end a…

Numerical Analysis · Mathematics 2021-01-19 Andrew Gloster

The study of binary pulsars enables tests of general relativity. Orbital motion in binary systems causes the apparent pulsar spin frequency to drift, reducing the sensitivity of periodicity searches. Acceleration searches are methods that…

Instrumentation and Methods for Astrophysics · Physics 2018-12-12 Sofia Dimoudi , Karel Adamek , Prabu Thiagaraj , Scott M. Ransom , Aris Karastergiou , Wesley Armour

Volumetric data structures typically prioritize data locality, focusing on efficient memory access patterns. This singular focus can neglect other critical performance factors, such as occupancy, communication, and kernel fusion. We…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-09-03 Massimiliano Meneghin , Ahmed H. Mahmoud

High-order gas-kinetic scheme (HGKS) has become a workable tool for the direct numerical simulation (DNS) of turbulence. In this paper, to accelerate the computation, HGKS is implemented with the graphical processing unit (GPU) using the…

Numerical Analysis · Mathematics 2023-01-25 Yuhang Wang , Guiyu Cao , Liang Pan

We present LLMQ, an end-to-end CUDA/C++ implementation for medium-sized language-model training, e.g. 3B to 32B parameters, on affordable, commodity GPUs. These devices are characterized by low memory availability and slow communication…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-18 Erik Schultheis , Dan Alistarh

The exact quantum dynamics of lattice models can be computationally intensive, especially when aiming for large system sizes and extended simulation times necessary to converge transport coefficients. By leveraging finite memory times to…

Chemical Physics · Physics 2024-11-14 Srijan Bhattacharyya , Thomas Sayer , Andrés Montoya-Castillo

We give an improved algorithm for learning a quantum Hamiltonian given copies of its Gibbs state, that can succeed at any temperature. Specifically, we improve over the work of Bakshi, Liu, Moitra, and Tang [BLMT24], by reducing the sample…

Quantum Physics · Physics 2024-07-08 Shyam Narayanan

We present extensive Monte Carlo simulations on a two-dimensional XY model with a modified form of interaction potential. Thermodynamic quantities other than energy, specific heat etc (such as magnetization, susceptibility, fourth order…

Statistical Mechanics · Physics 2013-05-24 Suman Sinha

This paper explores two condensed-space interior-point methods to efficiently solve large-scale nonlinear programs on graphics processing units (GPUs). The interior-point method solves a sequence of symmetric indefinite linear systems, or…

Optimization and Control · Mathematics 2025-08-15 François Pacaud , Sungho Shin , Alexis Montoison , Michel Schanen , Mihai Anitescu

Kernel matrix-vector multiplication (KMVM) is a foundational operation in machine learning and scientific computing. However, as KMVM tends to scale quadratically in both memory and time, applications are often limited by these…

Numerical Analysis · Mathematics 2025-02-25 Robert Hu , Siu Lun Chau , Dino Sejdinovic , Joan Alexis Glaunès

We discuss the application of a recently introduced numerical linked-cluster (NLC) algorithm to strongly correlated itinerant models. In particular, we present a study of thermodynamic observables: chemical potential, entropy, specific…

Strongly Correlated Electrons · Physics 2007-06-25 Marcos Rigol , Tyler Bryant , Rajiv R. P. Singh

In the next decade, the demands for computing in large scientific experiments are expected to grow tremendously. During the same time period, CPU performance increases will be limited. At the CERN Large Hadron Collider (LHC), these two…

Finetuning large language models (LLMs) is essential for task adaptation, yet today's serving stacks isolate inference and finetuning on separate GPU clusters -- wasting resources and under-utilizing hardware. We introduce FlexLLM, the…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-10-27 Gabriele Oliaro , Xupeng Miao , Xinhao Cheng , Vineeth Kada , Mengdi Wu , Ruohan Gao , Yingyi Huang , Remi Delacourt , April Yang , Yingcheng Wang , Colin Unger , Zhihao Jia
‹ Prev 1 8 9 10 Next ›