中文
相关论文

相关论文: Evaluation of performance portability frameworks f…

200 篇论文

With the proliferation of edge AI applications, satisfying user quality of experience (QoE) requirements, such as model inference latency, has become a first class objective, as these models operate in resource constrained settings and…

分布式、并行与集群计算 · 计算机科学 2025-12-15 Jianli Jin , Ziyang Lin , Qianli Dong , Yi Chen , Jayanth Srinivasa , Myungjin Lee , Zhaowei Tan , Fan Lai

Quantum-dot cellular automata (QCA) is a paradigm for low-power, general-purpose, classical computing designed to overcome the challenges facing CMOS in the extreme limits of scaling. A molecular implementation of QCA offers nanometer-scale…

量子物理 · 物理学 2022-08-31 Peizhong Cong , Enrique P. Blair

Fault tolerance is a long-term objective driving many companies and research organizations to compete in making current, imperfect quantum computers useful - Quantum Utility (QU). It looks promising to achieve this by leveraging software…

量子物理 · 物理学 2024-09-27 Markiian Tsymbalista , Ihor Katernyak

This paper presents an implementation of radio astronomy imaging algorithms on modern High Performance Computing (HPC) infrastructures, exploiting distributed memory parallelism and acceleration throughout multiple GPUs. Our code, called…

天体物理仪器与方法 · 物理学 2024-11-13 Emanuele De Rubeis , Giovanni Lacopo , Claudio Gheller , Luca Tornatore , Giuliano Taffoni

Massively parallel architectures offer the potential to significantly accelerate an application relative to their serial counterparts. However, not all applications exhibit an adequate level of data and/or task parallelism to exploit such…

计算物理 · 物理学 2018-08-08 Salvatore Cardamone , Jonathan R. Kimmitt , Hugh G. A. Burton , Alex J. W. Thom

In this work we use the GPU porting task for the operative Japanese weather prediction model "ASUCA" as an opportunity to examine productivity issues with OpenACC when applied to structured grid problems. We then propose "Hybrid Fortran",…

分布式、并行与集群计算 · 计算机科学 2017-12-11 Michel Müller , Takayuki Aoki

Cellular Automata(CA) is a discrete computing model which provides simple, flexible and efficient platform for simulating complicated systems and performing complex computation based on the neighborhoods information. CA consists of two…

分布式、并行与集群计算 · 计算机科学 2011-12-12 Debasis Das , Rajiv Misra

State-of-the-art numerical simulations of laser plasma by means of the Particle-in-Cell method are often extremely computationally intensive. Therefore there is a growing need for development of approaches for efficient utilization of…

Modeling plasma accelerators is a computationally challenging task and the quasi-static particle-in-cell algorithm is a method of choice in a wide range of situations. In this work, we present the first performance-portable, quasi-static,…

We present Porthos, the first tool that discovers porting bugs in performance-critical code. Porthos takes as input a program and the memory models of the source architecture for which the program has been developed and the target model to…

编程语言 · 计算机科学 2017-05-01 Hernán Ponce-de-León , Florian Furbach , Keijo Heljanko , Roland Meyer

We present a class of massively parallel processor architectures called invasive tightly coupled processor arrays (TCPAs). The presented processor class is a highly parameterizable template, which can be tailored before runtime to fulfill…

硬件体系结构 · 计算机科学 2014-05-14 Vahid Lari , Alexandru Tanase , Frank Hannig , Jürgen Teich

This paper presents the Parallel Coupler for Multimodel Simulations (PCMS), a new GPU accelerated generalized coupling framework for coupling simulation codes on leadership class supercomputers. PCMS includes distributed control and field…

分布式、并行与集群计算 · 计算机科学 2025-10-22 Jacob S. Merson , Cameron W. Smith , Mark S. Shephard , Fuad Hasan , Abhiyan Paudel , Angel Castillo-Crooke , Joyal Mathew , Mohammad Elahi

In this paper, we report on a preliminary investigation of the potential performance gain of programs implemented in field-programmable gate arrays (FPGAs) using a high-level language Chisel compared to ordinary high-level software…

性能 · 计算机科学 2022-05-27 Maja H. Kirkeby , Martin Schoeberl

Replica Exchange (RE) simulations have emerged as an important algorithmic tool for the molecular sciences. RE simulations involve the concurrent execution of independent simulations which infrequently interact and exchange information. The…

分布式、并行与集群计算 · 计算机科学 2016-01-22 Antons Treikalis , Andre Merzky , Haoyuan Chen , Tai-Sung Lee , Darrin M. York , Shantenu Jha

Robust PCA has drawn significant attention in the last decade due to its success in numerous application domains, ranging from bio-informatics, statistics, and machine learning to image and video processing in computer vision. Robust PCA…

最优化与控制 · 数学 2018-06-12 Shiqian Ma , Necdet Serhat Aybat

Parallel programmers face the often irreconcilable goals of programmability and performance. HPC systems use distributed memory for scalability, thereby sacrificing the programmability advantages of shared memory programming models.…

分布式、并行与集群计算 · 计算机科学 2013-01-21 Bharath Ramesh , Calvin J. Ribbens , Srinidhi Varadarajan

The Exascale Computing Project (ECP) is invested in co-design to assure that key applications are ready for exascale computing. Within ECP, the Co-design Center for Particle Applications (CoPA) is addressing challenges faced by…

This paper introduces a computer architecture, where part of the instruction set architecture (ISA) is implemented on small highly-integrated field-programmable gate arrays (FPGAs). Small FPGAs inside a general-purpose processor (CPU) can…

硬件体系结构 · 计算机科学 2022-08-23 Philippos Papaphilippou , Myrtle Shah

The world's largest particle accelerator, located at CERN, produces petabytes of data that need to be analysed efficiently, to study the fundamental structures of our universe. ROOT is an open-source C++ data analysis framework, developed…

分布式、并行与集群计算 · 计算机科学 2025-01-07 Jolly Chen , Monica Dessole , Ana Lucia Varbanescu

Molecular dynamics simulations are one of the methods in scientific computing that benefit from GPU acceleration. For those devices, SYCL is a promising API for writing portable codes. In this paper, we present the case study of "HAL's MD…

分布式、并行与集群计算 · 计算机科学 2024-06-07 Viktor Skoblin , Felix Höfling , Steffen Christgau
‹ 上一页 1 8 9 10 下一页 ›