English
Related papers

Related papers: A Parallel Computing Method for the Higher Order T…

200 papers

As parallel computing trends towards the exascale, scientific data produced by high-fidelity simulations are growing increasingly massive. For instance, a simulation on a three-dimensional spatial grid with 512 points per dimension that…

Numerical Analysis · Computer Science 2017-01-05 Woody Austin , Grey Ballard , Tamara G. Kolda

We introduce a general tensor model suitable for data analytic tasks for {\em heterogeneous} datasets, wherein there are joint low-rank structures within groups of observations, but also discriminative structures across different groups. To…

Machine Learning · Statistics 2022-10-04 Davoud Ataee Tarzanagh , George Michailidis

Training-free high-resolution (HR) image generation has garnered significant attention due to the high costs of training large diffusion models. Most existing methods begin by reconstructing the overall structure and then proceed to refine…

Computer Vision and Pattern Recognition · Computer Science 2024-12-10 Yuming Li , Peidong Jia , Daiwei Hong , Yueru Jia , Qi She , Rui Zhao , Ming Lu , Shanghang Zhang

Transformer-based models are becoming deeper and larger recently. For better scalability, an underlying training solution in industry is to split billions of parameters (tensors) into many tasks and then run them across homogeneous…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-01-23 Zhigang Wang , Xu Zhang , Ning Wang , Chuanfei Xu , Jie Nie , Zhiqiang Wei , Yu Gu , Ge Yu

We propose a new renormalization scheme of tensor networks made only of third order tensors. The isometry used for coarse-graining the network can be prepared at an $O(D^6)$ computational cost in any $d$ dimension ($d \ge 2$), where $D$ is…

High Energy Physics - Lattice · Physics 2019-12-06 Daisuke Kadoh , Katsumasa Nakayama

Interprocessor communication often dominates the runtime of large matrix computations. We present a parallel algorithm for computing QR decompositions whose bandwidth cost (communication volume) can be decreased at the cost of increasing…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-05-15 Grey Ballard , James Demmel , Laura Grigori , Mathias Jacquelin , Nicholas Knight

Parallelization techniques have become ubiquitous for accelerating inference and training of deep neural networks. Despite this, several operations are still performed in a sequential manner. For instance, the forward and backward passes…

Machine Learning · Computer Science 2023-10-30 Federico Danieli , Miguel Sarabia , Xavier Suau , Pau Rodríguez , Luca Zappella

We study the dynamical density matrix renormalization group (DDMRG) and time-dependent density matrix renormalization group (td-DMRG) algorithms in the ab initio context, to compute dynamical correlation functions of correlated systems. We…

Chemical Physics · Physics 2017-11-21 Enrico Ronca , Zhendong Li , Carlos A. Jimenez-Hoyos , Garnet Kin-Lic Chan

This work explores a distributed computing setting where $K$ nodes are assigned fractions (subtasks) of a computational task in order to perform the computation in parallel. In this setting, a well-known main bottleneck has been the…

Information Theory · Computer Science 2018-02-13 Emanuele Parrinello , Eleftherios Lampiris , Petros Elia

A distributed-memory parallelization strategy for the density matrix renormalization group is proposed for cases where correlation functions are required. This new strategy has substantial improvements with respect to previous works. A…

Strongly Correlated Electrons · Physics 2010-04-20 Julian Rincon , D. J. Garcia , K. Hallberg

We consider the problem of decomposing higher-order moment tensors, i.e., the sum of symmetric outer products of data vectors. Such a decomposition can be used to estimate the means in a Gaussian mixture model and for other applications in…

Numerical Analysis · Mathematics 2020-10-06 Samantha Sherman , Tamara G. Kolda

Tensor renormalization group (TRG) constitutes an important methodology for accurate simulations of strongly correlated lattice models. Facilitated by the automatic differentiation technique widely used in deep learning, we propose a…

Strongly Correlated Electrons · Physics 2020-07-07 Bin-Bin Chen , Yuan Gao , Yi-Bin Guo , Yuzhi Liu , Hui-Hai Zhao , Hai-Jun Liao , Lei Wang , Tao Xiang , Wei Li , Z. Y. Xie

We propose an improved tensor renormalization group (TRG) algorithm, the bond-weighted TRG (BTRG). In BTRG, we generalize the conventional TRG by introducing bond weights on the edges of the tensor network. We show that BTRG outperforms the…

Statistical Mechanics · Physics 2022-03-03 Daiki Adachi , Tsuyoshi Okubo , Synge Todo

High-speed long polynomial multiplication is important for applications in homomorphic encryption (HE) and lattice-based cryptosystems. This paper addresses low-latency hardware architectures for long polynomial modular multiplication using…

Hardware Architecture · Computer Science 2024-03-21 Weihang Tan , Sin-Wei Chiu , Antian Wang , Yingjie Lao , Keshab K. Parhi

We introduce a data distribution scheme for $\mathcal{H}$-matrices and a distributed-memory algorithm for $\mathcal{H}$-matrix-vector multiplication. Our data distribution scheme avoids an expensive $\Omega(P^2)$ scheduling procedure used…

Numerical Analysis · Mathematics 2020-09-23 Yingzhou Li , Jack Poulson , Lexing Ying

The Tucker decomposition generalizes the notion of Singular Value Decomposition (SVD) to tensors, the higher dimensional analogues of matrices. We study the problem of constructing the Tucker decomposition of sparse tensors on distributed…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-01-22 Venkatesan T. Chakaravarthy , Jee W. Choi , Douglas J. Joseph , Prakash Murali , Shivmaran S. Pandian , Yogish Sabharwal , Dheeraj Sreedhar

Numerical solution of partial differential equations on parallel computers using domain decomposition usually requires synchronization and communication among the processors. These operations often have a significant overhead in terms of…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-01-11 Soumyadip Ghosh , Jiacai Lu , Vijay Gupta , Gretar Tryggvason

Accurate and fast treatment of electron-electron interactions remains a central challenge in electronic structure theory because post-Hartree-Fock methods often suffered from the computational cost for 4-index electron repulsion integrals…

Chemical Physics · Physics 2025-12-05 Yueyang Zhang , Xuewei Xiong , Wei Wu , Peifeng Su

This paper presents \pandora, a novel parallel algorithm for efficiently constructing dendrograms for single-linkage hierarchical clustering, including \hdbscan. Traditional dendrogram construction methods from a minimum spanning tree…

Machine Learning · Computer Science 2025-04-29 Piyush Sao , Andrey Prokopenko , Damien Lebrun-Grandié

A parallel algorithm for the implementation of the recursive Green's function technique, which is extensively applied in the coherent scattering formalism, is developed. The algorithm performs a domain decomposition of the scattering region…

Mesoscale and Nanoscale Physics · Physics 2009-11-11 P. S. Drouvelis , P. Schmelcher , P. Bastian