中文
相关论文

相关论文: Analyzing HPC Support Tickets: Experience and Reco…

200 篇论文

Several scientific and industry applications require High Performance Computing (HPC) resources to process and/or simulate complex models. Not long ago, companies, research institutes, and universities used to acquire and maintain…

分布式、并行与集群计算 · 计算机科学 2015-07-21 Kiran Mantripragada , Leonardo P. Tizzei , Alecio P. D. Binotto , Marco A. S. Netto

High-performance computing (HPC) systems are a complex combination of software, processors, memory, networks, and storage systems characterized by frequent disruptive technological advances. Anomalous behavior has to be manually diagnosed…

分布式、并行与集群计算 · 计算机科学 2013-02-19 Charng-Da Lu

The next generation of High Energy Physics experiments are expected to generate exabytes of data---two orders of magnitude greater than the current generation. In order to reliably meet peak demands, facilities must either plan to provision…

High Performance Computing (HPC) is a highly demanded discipline in companies and institutions. However, as students and also afterwards as professors, we observed a lack of HPC related content in the engineering degrees at our university,…

分布式、并行与集群计算 · 计算机科学 2026-05-05 S. Catalán , R. Carratalá-Sáez , S. Iserte

The high-performance computing (HPC) community has adopted incentive structures to motivate reproducible research, with major conferences awarding badges to papers that meet reproducibility requirements. Yet, many papers do not meet such…

分布式、并行与集群计算 · 计算机科学 2025-09-01 Valérie Hayot-Sasson , Nathaniel Hudson , André Bauer , Maxime Gonthier , Ian Foster , Kyle Chard

The evolution of distributed architectures and programming paradigms for performance-oriented program development, challenge the state-of-the-art technology for performance tools. The area of high performance computing is rapidly expanding…

分布式、并行与集群计算 · 计算机科学 2010-06-15 Ajanta De Sarkar , Nandini Mukherjee

As scientific applications extend to the simulation of more and more complex systems, they involve an increasing number of abstraction levels, at each of which errors can emerge and across which they can propagate; tools for correctness…

软件工程 · 计算机科学 2011-01-18 Eloisa Bentivegna , Gabrielle Allen , Oleg Korobkin , Erik Schnetter

Decision support tools enable improved decision-making for challenging decision problems by empowering stakeholders to process, analyze, visualize, and otherwise make sense of a variety of key factors. Their intentional design is a critical…

计算机与社会 · 计算机科学 2021-11-12 Narges Ahani , Andrew C. Trapp

The rise of AI and the economic dominance of cloud computing have created a new nexus of innovation for high performance computing (HPC), which has a long history of driving scientific discovery. In addition to performance needs, scientific…

分布式、并行与集群计算 · 计算机科学 2025-06-10 Vanessa Sochat , Daniel Milroy , Abhik Sarkar , Aniruddha Marathe , Tapasya Patki

Monitoring and Managing High Performance Computing (HPC) systems and environments generate an ever growing amount of data. Making sense of this data and generating a platform where the data can be visualized for system administrators and…

数据库 · 计算机科学 2019-02-12 Rebecca Wild , Matthew Hubbell , Jeremy Kepner

Over the past decade, high performance computational (HPC) clusters have become mainstream in academic and industrial settings as accessible means of computation. Throughout their proliferation, HPC security has been a secondary concern to…

密码学与安全 · 计算机科学 2007-05-23 Dmitry Mogilevsky , Adam Lee , William Yurcik

High-Performance Computing (HPC) systems need to be constantly monitored to ensure their stability. The monitoring systems collect a tremendous amount of data about different parameters or Key Performance Indicators (KPIs), such as resource…

人工智能 · 计算机科学 2023-12-12 Mohamed Soliman Halawa , Rebeca P. Díaz-Redondo , Ana Fernández-Vilas

We recognize the emergence of a statistical computing community focused on working with large computing platforms and producing software and applications that exemplify high-performance statistical computing (HPSC). The statistical…

分布式、并行与集群计算 · 计算机科学 2025-11-03 Sameh Abdulah , Mary Lai O. Salvana , Ying Sun , David E. Keyes , Marc G. Genton

Over time, software systems suffer gradual quality decay and therefore costs can rise if organizations fail to take proactive countermeasures. Quality control is the first step to avoiding this cost trap. Continuous quality assessments help…

High-performance computing (HPC) requires resilience techniques such as checkpointing in order to tolerate failures in supercomputers. As the number of nodes and memory in supercomputers keeps on increasing, the size of checkpoint data also…

分布式、并行与集群计算 · 计算机科学 2019-06-13 Kai Keller , Leonardo Bautista Gomez

Nowadays, improving the energy efficiency of high-performance computing (HPC) systems is one of the main drivers in scientific and technological research. As large-scale HPC systems require some fault-tolerant method, the opportunities to…

分布式、并行与集群计算 · 计算机科学 2023-11-15 Marina Moran , Javier Balladini , Dolores Rexachs , Enzo Rucci

Background: It has long been suggested that user feedback, typically written in natural language by end-users, can help issue detection. However, for large-scale online service systems that receive a tremendous amount of feedback, it…

软件工程 · 计算机科学 2025-08-04 Shuyao Jiang , Jiazhen Gu , Wujie Zheng , Yangfan Zhou , Michael R. Lyu

Support teams of high-performance computing (HPC) systems often find themselves between a rock and a hard place: on one hand, they understandably administrate these large systems in a conservative way, but on the other hand, they try to…

分布式、并行与集群计算 · 计算机科学 2015-07-28 Ludovic Courtès , Ricardo Wurmus

High-performance computing (HPC) centers consume substantial power, incurring environmental and operational costs. This review assesses how artificial intelligence (AI), including machine learning (ML) and optimization, improves the…

分布式、并行与集群计算 · 计算机科学 2026-02-03 Pierrick Pochelu , Hyacinthe Cartiaux , Julien Schleich

This paper reports on the design and implementation of the HPC performance monitoring system deployed to continuously monitor performance metrics of all jobs on the HPC systems at the Max Planck Computing and Data Facility (MPCDF). Thereby…

分布式、并行与集群计算 · 计算机科学 2019-09-27 Luka Stanisic , Klaus Reuter