English
Related papers

Related papers: Multi GPU Performance of Conjugate Gradient Algori…

200 papers

Edge computing's growing prominence, due to its ability to reduce communication latency and enable real-time processing, is promoting the rise of high-performance, heterogeneous System-on-Chip solutions. While current approaches often…

Artificial Intelligence · Computer Science 2024-09-24 Rakshith Jayanth , Neelesh Gupta , Viktor Prasanna

Availability of high performance computing infrastructures such as clusters of GPUs and CPUs have fueled the growth of distributed learning systems. Deep Learning frameworks express neural nets as DAGs and execute these DAGs on computation…

Distributed, Parallel, and Cluster Computing · Computer Science 2018-02-21 Amith R Mamidala

We describe GPU implementations of the matrix recommender algorithms CCD++ and ALS. We compare the processing time and predictive ability of the GPU implementations with existing multi-core versions of the same algorithms. Results on the…

Information Retrieval · Computer Science 2015-11-10 André Valente Rodrigues , Alípio Jorge , Inês Dutra

Graphics Processing Units (GPUs) consisting of Streaming Multiprocessors (SMs) achieve high throughput by running a large number of threads and context switching among them to hide execution latencies. The number of thread blocks, and hence…

Hardware Architecture · Computer Science 2015-06-08 Vishwesh Jatala , Jayvant Anantpur , Amey Karkare

We design, implement, and evaluate GPU-based algorithms for the maximum cardinality matching problem in bipartite graphs. Such algorithms have a variety of applications in computer science, scientific computing, bioinformatics, and other…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-03-07 Mehmet Deveci , Kamer Kaya , Bora Ucar , Umit V. Catalyurek

Load-balancing among the threads of a GPU for graph analytics workloads is difficult because of the irregular nature of graph applications and the high variability in vertex degrees, particularly in power-law graphs. We describe a novel…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-02-28 Vishwesh Jatala , Loc Hoang , Roshan Dathathri , Gurbinder Gill , V Krishna Nandivada , Keshav Pingali

Modern heterogeneous supercomputing systems are comprised of CPUs, GPUs, and high-speed network interconnects. Communication libraries supporting efficient data transfers involving memory buffers from the GPU memory typically require the…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-29 Naveen Namashivayam , Krishna Kandalla , James B White , Larry Kaplan , Mark Pagel

We investigate the performance of Opticks, a NVIDIA OptiX API 7.5 GPU-accelerated photon propagation tool compared with a single-threaded Geant4 simulation. We compare the simulations using an improved model of the NEXT-CRAB-0 gaseous time…

Instrumentation and Detectors · Physics 2025-11-20 NEXT Collaboration , I. Parmaksiz , K. Mistry , E. Church , C. Adams , J. Asaadi , J. Baeza-Rubio , K. Bailey , N. Byrnes , B. J. P. Jones , I. A. Moya , K. E. Navarro , D. R. Nygren , P. Oyedele , L. Rogers , F. Samaniego , K. Stogsdill , H. Almazán , V. Álvarez , B. Aparicio , A. I. Aranburu , L. Arazi , I. J. Arnquist , F. Auria-Luna , S. Ayet , C. D. R. Azevedo , F. Ballester , M. del Barrio-Torregrosa , A. Bayo , J. M. Benlloch-Rodríguez , F. I. G. M. Borges , A. Brodolin , S. Cárcel , A. Castillo , L. Cid , C. A. N. Conde , T. Contreras , F. P. Cossío , R. Coupe , E. Dey , G. Díaz , C. Echevarria , M. Elorza , J. Escada , R. Esteve , R. Felkai , L. M. P. Fernandes , P. Ferrario , A. L. Ferreira , F. W. Foss , Z. Freixa , J. García-Barrena , J. J. Gómez-Cadenas , J. W. R. Grocott , R. Guenette , J. Hauptman , C. A. O. Henriques , J. A. Hernando Morata , P. Herrero-Gómez , V. Herrero , C. Hervés Carrete , Y. Ifergan , F. Kellerer , L. Larizgoitia , A. Larumbe , P. Lebrun , F. Lopez , N. López-March , R. Madigan , R. D. P. Mano , A. P. Marques , J. Martín-Albo , G. Martínez-Lema , M. Martínez-Vara , R. L. Miller , J. Molina-Canteras , F. Monrabal , C. M. B. Monteiro , F. J. Mora , P. Novella , A. Nuñez , E. Oblak , J. Palacio , B. Palmeiro , A. Para , A. Pazos , J. Pelegrin , M. Pérez Maneiro , M. Querol , J. Renner , I. Rivilla , C. Rogero , B. Romeo , C. Romo-Luque , V. San Nacienciano , F. P. Santos , J. M. F. dos Santos , M. Seemann , I. Shomroni , P. A. O. C. Silva , A. Simón , S. R. Soleti , M. Sorel , J. Soto-Oton , J. M. R. Teixeira , S. Teruel-Pardo , J. F. Toledo , C. Tonnelé , S. Torelli , J. Torrent , A. Trettin , A. Usón , P. R. G. Valle , J. F. C. A. Veloso , J. Waiton , A. Yubero-Navarro

The continued growth in the processing power of FPGAs coupled with high bandwidth memories (HBM), makes systems like the Xilinx U280 credible platforms for linear solvers which often dominate the run time of scientific and engineering…

Hardware Architecture · Computer Science 2023-01-02 Linghao Song , Licheng Guo , Suhail Basalama , Yuze Chi , Robert F. Lucas , Jason Cong

As an increasing number of leadership-class systems embrace GPU accelerators in the race towards exascale, efficient communication of GPU data is becoming one of the most critical components of high-performance computing. For developers of…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-03-23 Jaemin Choi , Zane Fink , Sam White , Nitin Bhat , David F. Richards , Laxmikant V. Kale

Graph coloring has been broadly used to discover concurrency in parallel computing. To speedup graph coloring for large-scale datasets, parallel algorithms have been proposed to leverage modern GPUs. Existing GPU implementations either have…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-01-22 Xuhao Chen , Pingfan Li , Jianbin Fang , Tao Tang , Zhiying Wang , Canqun Yang

In classification problems with large output spaces (up to millions of labels), the last layer can require an enormous amount of memory. Using sparse connectivity would drastically reduce the memory requirements, but as we show below, it…

Machine Learning · Computer Science 2023-11-08 Erik Schultheis , Rohit Babbar

The molecular dynamics simulation package GROMACS runs efficiently on a wide variety of hardware from commodity workstations to high performance computing clusters. Hardware features are well exploited with a combination of SIMD,…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-03-14 Carsten Kutzner , Szilárd Páll , Martin Fechner , Ansgar Esztermann , Bert L. de Groot , Helmut Grubmüller

Training and deploying large-scale machine learning models is time-consuming, requires significant distributed computing infrastructures, and incurs high operational costs. Our analysis, grounded in real-world large model training on…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-06-12 Samuel Hsia , Alicia Golden , Bilge Acun , Newsha Ardalani , Zachary DeVito , Gu-Yeon Wei , David Brooks , Carole-Jean Wu

Correlation Plenoptic Imaging (CPI) is a novel technological imaging modality enabling to overcome drawbacks of standard plenoptic devices, while preserving their advantages. However, a major challenge in view of real-time application of…

Modern large-scale computing systems (data centers, supercomputers, cloud and edge setups and high-end cyber-physical systems) employ heterogeneous architectures that consist of multicore CPUs, general-purpose many-core GPUs, and…

Discrete optimization is a central problem in artificial intelligence. The optimization of the aggregated cost of a network of cost functions arises in a variety of problems including (W)CSP, DCOP, as well as optimization in stochastic…

Artificial Intelligence · Computer Science 2018-01-12 Ferdinando Fioretto , Enrico Pontelli , William Yeoh , Rina Dechter

Large scale graph optimization problems arise in many fields. This paper presents an extensible, high performance framework (named OpenGraphGym-MG) that uses deep reinforcement learning and graph embedding to solve large graph optimization…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-06-25 Weijian Zheng , Dali Wang , Fengguang Song

This paper investigates a new method for improving the learning algorithm of Mixture of Experts (ME) model using a hybrid of Modified Cuckoo Search (MCS) and Conjugate Gradient (CG) as a second order optimization technique. The CG technique…

Artificial Intelligence · Computer Science 2012-02-20 Hamid Salimi , Davar Giveki , Mohammad Ali Soltanshahi , Javad Hatami

Betweenness centrality (BC) is an important graph analytical application for large-scale graphs. While there are many efforts for parallelizing betweenness centrality algorithms on multi-core CPUs and many-core GPUs, in this work, we…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-08-14 Ashirbad Mishra , Sathish Vadhiyar , Rupesh Nasre , Keshav Pingali