English
Related papers

Related papers: Deploying a Top-100 Supercomputer for Large Parall…

200 papers

The Superfacility model is designed to leverage HPC for experimental science. It is more than simply a model of connected experiment, network, and HPC facilities; it encompasses the full ecosystem of infrastructure, software, tools, and…

High fidelity Computational Fluid Dynamics simulations are generally associated with large computing requirements, which are progressively acute with each new generation of supercomputers. However, significant research efforts are required…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-07-07 R. Borrell , D. Dosimont , M. Garcia-Gasulla , G. Houzeaux , O. Lehmkuhl , V. Mehta , H. Owen , M. Vazquez , G. Oyarzun

Many of the most performant deep learning models today in fields like language and image understanding are fine-tuned models that contain billions of parameters. In anticipation of workloads that involve serving many of such large models to…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-06-27 Daniel Zou , Xinchen Jin , Xueyang Yu , Hao Zhang , James Demmel

The Center for Exascale Monte Carlo Neutron Transport is developing Monte Carlo / Dynamic Code (MC/DC) as a portable Monte Carlo neutron transport package for rapid numerical methods exploration on CPU- and GPU-based high-performance…

Computational Physics · Physics 2025-05-30 Joanna Piper Morgan , Braxton Cuneo , Ilham Variansyah , Kyle E. Niemeyer

Neural Processing Units (NPUs) are key to enabling efficient AI inference in resource-constrained edge environments. While peak tera operations per second (TOPS) is often used to gauge performance, it poorly reflects real-world performance…

Hardware Architecture · Computer Science 2025-09-19 Lennart Bamberg , Filippo Minnella , Roberto Bosio , Fabrizio Ottati , Yuebin Wang , Jongmin Lee , Luciano Lavagno , Adam Fuks

Aurora is Argonne National Laboratory's pioneering Exascale supercomputer, designed to accelerate scientific discovery with cutting-edge architectural innovations. Key new technologies include the Intel(TM) Xeon(TM) Data Center GPU Max…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-09 William E. Allcock , Benjamin S. Allen , James Anchell , Victor Anisimov , Thomas Applencourt , Abhishek Bagusetty , Ramesh Balakrishnan , Riccardo Balin , Solomon Bekele , Colleen Bertoni , Cyrus Blackworth , Renzo Bustamante , Kevin Canada , John Carrier , Christopher Chan-nui , Lance C. Cheney , Taylor Childers , Paul Coffman , Susan Coghlan , Tanima Dey , Michael D'Mello , Ashok Emani , Murali Emani , Kyle G. Felker , Sam Foreman , Olivier Franza , Longfei Gao , Marta García , María Garzarán , Balazs Gerofi , Yasaman Ghadar , Subrata Goswami , Neha Gupta , Kevin Harms , Väinö Hatanpää , Brian Holland , Carissa Holohan , Brian Homerding , Khalid Hossain , Xue Hu , Louise Huot , Huda Ibeid , Joseph A. Insley , Sai Jayanthi , Hong Jiang , Wei Jiang , Xiao-Yong Jin , Jeongnim Kim , Christopher Knight , Panagiotis Kourdis , Kalyan Kumaran , JaeHyuk Kwack , Janghaeng Lee , Ti Leggett , Ben Lenard , Chris Lewis , Nevin Liber , Johann Lombardi , Raymond M. Loy , Ye Luo , Bethany Lusch , Nilakantan Mahadevan , Beth Markey , Victor A. Mateevitsi , Gordon McPheeters , Ryan Milner , Jerome Mitchell , Vitali A. Morozov , Servesh Muralidharan , Tom Musta , Mrigendra Nagar , Vikram Narayana , Marieme Ngom , Anthony-Trung Nguyen , Nathan Nichols , Aditya Nishtala , James C. Osborn , Michael E. Papka , Scott Parker , Saumil S. Patel , Julia Piotrowska , Adrian C. Pope , Sucheta Raghunanda , Esteban Rangel , Paul M. Rich , Katherine M. Riley , Silvio Rizzi , Kris Rowe , Varuni Sastry , Adam Scovel , Filippo Simini , Haritha Siddabathuni Som , Patrick Steinbrecher , Rick Stevens , Xinmin Tian , Peter Upton , Thomas Uram , Archit K. Vasan , Álvaro Vázquez-Mayagoitia , Kaushik Velusamy , Brice Videau , Venkatram Vishwanath , Brian Whitney , Timothy J. Williams , Michael Woodacre , Sam Zeltner , Chuanjun Zhang , Gengbin Zheng , Huihuo Zheng

Network traffic is difficult to monitor and analyze, especially in high-bandwidth networks. Performance analysis, in particular, presents extreme complexity and scalability challenges. GPU (Graphics Processing Unit) technology has been…

Networking and Internet Architecture · Computer Science 2011-08-09 Wenji Wu , Phil DeMar , Don Holmgren , Amitoj Singh , Ruth Pordes

In emerging scientific computing environments, matrix computations of increasing size and complexity are increasingly becoming prevalent. However, contemporary matrix language implementations are insufficient in their support for efficient…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-12-11 Jay Hwan Lee , Yeonsoo Kim , Younghyun Ryu , Wasuwee Sodsong , Hyunjun Jeon , Jinsik Park , Bernd Burgstaller , Bernhard Scholz

Heterogeneous systems have become one of the most common architectures today, thanks to their excellent performance and energy consumption. However, due to their heterogeneity they are very complex to program and even more to achieve…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-10-26 Raúl Nozal , Jose Luis Bosque , Ramón Beivide

A heterogeneous memory has a single address space with fast access to some addresses (a fast tier of DRAM) and slow access to other addresses (a capacity tier of CXL-attached memory or NVM). A tiered memory system aims to maximize the…

Emerging Technologies · Computer Science 2025-10-28 Rohan Kadekodi , Haoran Peng , Gilbert Bernstein , Michael D. Ernst , Baris Kasikci

Designing and implementing efficient, provably correct parallel neural network processing is challenging. Existing high-level parallel abstractions like MapReduce are insufficiently expressive while low-level tools like MPI and Pthreads…

Machine Learning · Computer Science 2016-06-21 Maohua Zhu , Liu Liu , Chao Wang , Yuan Xie

Access to parallel and distributed computation has enabled researchers and developers to improve algorithms and performance in many applications. Recent research has focused on next generation special purpose systems with multiple kinds of…

Machine Learning · Computer Science 2019-06-11 Tegg Taekyong Sung , Valliappa Chockalingam , Alex Yahja , Bo Ryu

This paper describes JANUS, a modular massively parallel and reconfigurable FPGA-based computing system. Each JANUS module has a computational core and a host. The computational core is a 4x4 array of FPGA-based processing elements with…

We propose an approach to utilize idle computational resources of supercomputers. The idea is to maintain an additional queue of low-priority non-parallel jobs and execute them in containers, using container migration tools to break the…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-09-04 Julia Dubenskaya , Stanislav Polyakov

A new pre-exascale computer cluster has been designed to foster scientific progress and competitive innovation across European research systems, it is called LEONARDO. This paper describes the general architecture of the system and focuses…

Distributed, Parallel, and Cluster Computing · Computer Science 2023-08-01 Matteo Turisini , Giorgio Amati , Mirko Cestari

We present Synkhronos, an extension to Theano for multi-GPU computations leveraging data parallelism. Our framework provides automated execution and synchronization across devices, allowing users to continue to write serial programs without…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-10-13 Adam Stooke , Pieter Abbeel

Distributed training frameworks, like TensorFlow, have been proposed as a means to reduce the training time of deep learning models by using a cluster of GPU servers. While such speedups are often desirable---e.g., for rapidly evaluating…

Performance · Computer Science 2019-05-07 Shijian Li , Robert J. Walls , Lijie Xu , Tian Guo

In this paper, we describe the architecture and performance of the GraCCA system, a Graphic-Card Cluster for Astrophysics simulations. It consists of 16 nodes, with each node equipped with 2 modern graphic cards, the NVIDIA GeForce 8800…

Astrophysics · Physics 2008-11-26 Hsi-Yu Schive , Chia-Hung Chien , Shing-Kwong Wong , Yu-Chih Tsai , Tzihong Chiueh

TriCloudEdge is a scalable three-tier cloud continuum that integrates far-edge devices, intermediate edge nodes, and central cloud services, working in parallel as a unified solution. At the far edge, ultra-low-cost microcontrollers can…

Networking and Internet Architecture · Computer Science 2026-02-17 George Violettas , Lefteris Mamatas

The evolution of high-performance computing is associated with the growth of energy consumption. Performance of cluster computes (is increased via rising in performance and the number of used processors, GPUs, and coprocessors. An increment…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-12-23 E. A. Kiselev , P. N. Telegin , B. M. Shabanov
‹ Prev 1 4 5 6 7 8 10 Next ›