English
Related papers

Related papers: Acceleration of the tree method with SIMD instruct…

200 papers

We describe a method for fast approximation of sparse coding. The input space is subdivided by a binary decision tree, and we simultaneously learn a dictionary and assignment of allowed dictionary elements for each leaf of the tree. We…

Computer Vision and Pattern Recognition · Computer Science 2015-06-09 Arthur Szlam , Karol Gregor , Yann LeCun

Simulation of quantum systems is challenging due to the exponential size of the state space. Tensor networks provide a systematically improvable approximation for quantum states. 2D tensor networks such as Projected Entangled Pair States…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-09-04 Yuchen Pang , Tianyi Hao , Annika Dugad , Yiqing Zhou , Edgar Solomonik

On modern architectures, the performance of 32-bit operations is often at least twice as fast as the performance of 64-bit operations. By using a combination of 32-bit and 64-bit floating point arithmetic, the performance of many dense and…

Mathematical Software · Computer Science 2015-05-13 Marc Baboulin , Alfredo Buttari , Jack Dongarra , Jakub Kurzak , Julie Langou , Julien Langou , Piotr Luszczek , Stanimire Tomov

Ground-state auxiliary-field quantum Monte Carlo (AFQMC) methods have become key numerical tools for studying quantum phases and phase transitions in interacting many-fermion systems. Despite the broad applicability, the efficiency of these…

Strongly Correlated Electrons · Physics 2026-03-23 Hao Du , Yuan-Yao He

When simulating a lattice system near its critical temperature, local algorithms for modeling the system's evolution can introduce very large autocorrelation times into sampled data. This critical slowing down places restrictions on the…

High Energy Physics - Lattice · Physics 2023-03-01 Tristan Protzman , Joel Giedt

With high computation power and memory bandwidth, graphics processing units (GPUs) lend themselves to accelerate data-intensive analytics, especially when such applications fit the single instruction multiple data (SIMD) model. However,…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-12-12 Hang Liu , H. Howie Huang

This paper proposes an accelerated proximal point method for maximally monotone operators. The proof is computer-assisted via the performance estimation problem approach. The proximal point method includes various well-known convex…

Optimization and Control · Mathematics 2021-03-25 Donghwan Kim

Upcoming Large Scale Structure surveys aim to achieve an unprecedented level of precision in measuring galaxy clustering. However, accurately modeling these statistics may require theoretical templates that go beyond second-order…

Cosmology and Nongalactic Astrophysics · Physics 2023-11-21 M. Icaza-Lizaola , Yong-Seon Song , Minji Oh , Yi Zheng

Quantum dynamics can be simulated on a quantum computer by exponentiating elementary terms from the Hamiltonian in a sequential manner. However, such an implementation of Trotter steps has gate complexity depending on the total Hamiltonian…

Quantum Physics · Physics 2023-05-15 Guang Hao Low , Yuan Su , Yu Tong , Minh C. Tran

Training LLMs presents significant memory challenges due to growing size of data, weights, and optimizer states. Techniques such as data and model parallelism, gradient checkpointing, and offloading strategies address this issue but are…

Machine Learning · Computer Science 2024-10-22 Arijit Das

We describe several additions to the ENIGMA system that guides clause selection in the E automated theorem prover. First, we significantly speed up its neural guidance by adding server-based GPU evaluation. The second addition is motivated…

Artificial Intelligence · Computer Science 2021-07-16 Zarathustra Goertzel , Karel Chvalovský , Jan Jakubův , Miroslav Olšák , Josef Urban

We present an implementation of the hierarchical tree algorithm on the individual timestep algorithm (the Hermite scheme) for collisional $N$-body simulations, running on GRAPE-9 system, a special-purpose hardware accelerator for…

Instrumentation and Methods for Astrophysics · Physics 2016-03-23 Toshiyuki Fukushige , Atsushi Kawai

In current computer architectures, data movement (from die to network) is by far the most energy consuming part of an algorithm (10pJ/word on-die to 10,000pJ/word on the network). To increase memory locality at the hardware level and reduce…

Computational Physics · Physics 2018-01-17 H. Vincenti , R. Lehe , R. Sasanka , J-L. Vay

With the widespread use of self-consistent field methods, including Hartree-Fock and Density Functional Theory, the implications of accelerating these methods are immense. To this end, we develop a tensor hypercontraction construction with…

Chemical Physics · Physics 2025-08-27 Andreas Erbs Hillers-Bendtsen , Todd J. Martínez

In this paper we explore acceleration techniques for large scale nonconvex optimization problems with special focuses on deep neural networks. The extrapolation scheme is a classical approach for accelerating stochastic gradient descent for…

Machine Learning · Statistics 2018-05-18 Guangzeng Xie , Yitan Wang , Shuchang Zhou , Zhihua Zhang

We present a method for efficient differentiable simulation of articulated bodies. This enables integration of articulated body dynamics into deep learning frameworks, and gradient-based optimization of neural networks that operate on…

Machine Learning · Computer Science 2021-09-17 Yi-Ling Qiao , Junbang Liang , Vladlen Koltun , Ming C. Lin

A Hessian based optimal control method is presented in Liouville space to mitigate previously undesirable polynomial scaling of computation time. This new method, an improvement to the state-of-the-art Newton-Raphson GRAPE method, is…

Optimization and Control · Mathematics 2022-08-19 David L. Goodwin , Mads Sloth Vinding

Among the algorithms that are likely to play a major role in future exascale computing, the fast multipole method (FMM) appears as a rising star. Our previous recent work showed scaling of an FMM on GPU clusters, with problem sizes in the…

Numerical Analysis · Computer Science 2012-10-30 Rio Yokota , Lorena Barba

We present Sapporo, a library for performing high-precision gravitational N-body simulations on NVIDIA Graphical Processing Units (GPUs). Our library mimics the GRAPE-6 library, and N-body codes currently running on GRAPE-6 can switch to…

Instrumentation and Methods for Astrophysics · Physics 2015-05-13 Evghenii Gaburov , Stefan Harfst , Simon Portegies Zwart

In this paper we consider the time complexity of computing the sum and product of two $n$-bit numbers within the tile self-assembly model. The (abstract) tile assembly model is a mathematical model of self-assembly in which system…

Data Structures and Algorithms · Computer Science 2013-08-06 Alexandra Keenan , Robert Schweller , Michael Sherman , Xingsi Zhong