中文
相关论文

相关论文: Performance of the Cray T3D and Emerging Architect…

200 篇论文

The relaxed semantics and rich functionality of one-sided communication primitives of MPI-3 makes MPI an attractive candidate for the implementation of PGAS models. However, the performance of such implementation suffers from the fact, that…

分布式、并行与集群计算 · 计算机科学 2016-03-08 Huan Zhou , Kamran Idrees , José Gracia

This paper investigates the multi-GPU performance of a 3D buoyancy driven cavity solver using MPI and OpenACC directives on different platforms. The paper shows that decomposing the total problem in different dimensions affects the strong…

分布式、并行与集群计算 · 计算机科学 2021-06-10 Weicheng Xue , Christopher J. Roy

The Unified Model (UM) code supports simulation of weather, climate and earth system processes. It is primarily developed by the UK Met Office, but in recent years a wider community of users and developers have grown around the code. Here…

计算物理 · 物理学 2015-11-13 Karthee Sivalingam , Grenville Lister , Bryan Lawrence

Network congestion in high-speed interconnects is a major source of application run time performance variation. Recent years have witnessed a surge of interest from both academia and industry in the development of novel approaches for…

分布式、并行与集群计算 · 计算机科学 2019-07-12 Saurabh Jha , Archit Patke , Jim Brandt , Ann Gentile , Mike Showerman , Eric Roman , Zbigniew T. Kalbarczyk , William T. Kramer , Ravishankar K. Iyer

Asynchronous tasks, when created with over-decomposition, enable automatic computation-communication overlap which can substantially improve performance and scalability. This is not only applicable to traditional CPU-based systems, but also…

分布式、并行与集群计算 · 计算机科学 2022-03-23 Jaemin Choi , David F. Richards , Laxmikant V. Kale

Concurrent priority queues are widely used in important workloads, such as graph applications and discrete event simulations. However, designing scalable concurrent priority queues for NUMA architectures is challenging. Even though several…

分布式、并行与集群计算 · 计算机科学 2024-06-12 Christina Giannoula , Foteini Strati , Dimitrios Siakavaras , Georgios Goumas , Nectarios Koziris

Two-phase I/O is a well-known strategy for implementing collective MPI-IO functions. It redistributes I/O requests among the calling processes into a form that minimizes the file access costs. As modern parallel computers continue to grow…

分布式、并行与集群计算 · 计算机科学 2020-06-23 Qiao Kang , Sunwoo Lee , Kai-yuan Hou , Robert Ross , Ankit Agrawal , Alok Choudhary , Wei-keng Liao

N-body algorithms for long-range unscreened interactions like gravity belong to a class of highly irregular problems whose optimal solution is a challenging task for present-day massively parallel computers. In this paper we describe a…

计算物理 · 物理学 2009-10-30 U. Becciani , R. Ansaloni , V. Antonuccio-Delogu , G. Erbacci , M. Gambera , A. Pagliaro , -

3D integration has the potential to improve the scalability and performance of Chip Multiprocessors (CMP). A closed form analytical solution for optimizing 3D CMP cache hierarchy is developed. It allows optimal partitioning of the cache…

硬件体系结构 · 计算机科学 2013-11-08 Leonid Yavits , Amir Morad , Ran Ginosar

MPI is the most widely used data transfer and communication model in High Performance Computing. The latest version of the standard, MPI-3, allows skilled programmers to exploit all hardware capabilities of the latest and future…

分布式、并行与集群计算 · 计算机科学 2016-09-30 Huan Zhou , Jose Gracia

The explosively growing communication traffic in datacenters imposes increasingly stringent performance requirements on the underlying networks. Over the last years, researchers have developed innovative optical switching technologies that…

网络与互联网体系结构 · 计算机科学 2024-06-21 Johannes Zerwas , Chen Griner , Stefan Schmid , Chen Avin

High-performance computing (HPC) systems frequently experience congestion leading to significant application performance variation. However, the impact of congestion on application runtime differs from application to application depending…

分布式、并行与集群计算 · 计算机科学 2021-02-05 Archit Patke , Saurabh Jha , Haoran Qiu , Jim Brandt , Ann Gentile , Joe Greenseid , Zbigniew Kalbarczyk , Ravishankar Iyer

The advent of multi-/many-core processors in clusters advocates hybrid parallel programming, which combines Message Passing Interface (MPI) for inter-node parallelism with a shared memory model for on-node parallelism. Compared to the…

分布式、并行与集群计算 · 计算机科学 2020-07-15 Huan Zhou , Jose Gracia , Ralf Schneider

The performance of large-scale computing systems often critically depends on high-performance communication networks. Dynamically reconfigurable topologies, e.g., based on optical circuit switches, are emerging as an innovative new…

网络与互联网体系结构 · 计算机科学 2022-12-29 Vamsi Addanki , Chen Avin , Stefan Schmid

In Network on Chip (NoC) rooted system, energy consumption is affected by task scheduling and allocation schemes which affect the performance of the system. In this paper we test the pre-existing proposed algorithms and introduced a new…

其他计算机科学 · 计算机科学 2014-05-02 Vaibhav Jha , Mohit Jha , GK Sharma

Learned indexes, which use machine learning models to replace traditional index structures, have shown promising results in recent studies. However, existing learned indexes exhibit a performance gap between synthetic and real-world…

数据库 · 计算机科学 2022-05-20 Jiaoyi Zhang , Yihan Gao

Complex applications and workflows needs are often exclusively expressed in terms of computational resources on HPC systems. In many cases, other resources like storage or network are not allocatable and are shared across the entire HPC…

分布式、并行与集群计算 · 计算机科学 2020-01-10 François Tessier , Maxime Martinasso , Matteo Chesi , Mark Klein , Miguel Gila

Production-quality parallel applications are often a mixture of diverse operations, such as computation- and communication-intensive, regular and irregular, tightly coupled and loosely linked operations. In conventional construction of…

分布式、并行与集群计算 · 计算机科学 2017-08-07 Ivy Bo Peng , Roberto Gioiosa , Gokcen Kestor , Erwin Laure , Stefano Markidis

Deep neural networks (DNNs) offer plenty of challenges in executing efficient computation at edge nodes, primarily due to the huge hardware resource demands. The article proposes HYDRA, hybrid data multiplexing, and runtime layer…

硬件体系结构 · 计算机科学 2026-03-31 Sonu Kumar , Komal Gupta , Gopal Raut , Mukul Lokhande , Santosh Kumar Vishvakarma

Memory tiering has received wide adoption in recent years as an effective solution to address the increasing memory demands of memory-intensive workloads. However, existing tiered memory systems often fail to meet service-level objectives…

操作系统 · 计算机科学 2024-12-13 Jiaheng Lu , Yiwen Zhang , Hasan Al Maruf , Minseo Park , Yunxuan Tang , Fan Lai , Mosharaf Chowdhury
‹ 上一页 1 2 3 10 下一页 ›