English
Related papers

Related papers: DD-$\alpha$AMG on QPACE 3

200 papers

Over the last ten years, graphics processors have become the de facto accelerator for data-parallel tasks in various branches of high-performance computing, including machine learning and computational sciences. However, with the recent…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-05-28 Johannes Pekkilä , Oskar Lappi , Fredrik Robertsén , Maarit J. Korpi-Lagg

Transposed Convolutions (TCONV) enable the up-scaling mechanism within generative Artificial Intelligence (AI) models. However, the predominant Input-Oriented Mapping (IOM) method for implementing TCONV has complex output mapping,…

Hardware Architecture · Computer Science 2025-07-11 Jude Haris , José Cano

We propose a unified Transformer-based architecture for wireless signal processing tasks, offering a low-latency, task-adaptive alternative to conventional receiver pipelines. Unlike traditional modular designs, our model integrates channel…

Signal Processing · Electrical Eng. & Systems 2025-09-11 Yuto Kawai , Rajeev Koodli

Quantum computers represent a transformative frontier in computational technology, promising exponential speedups beyond classical computing limits. IBM Quantum has led significant advancements in both hardware and software, providing…

Quantum Physics · Physics 2025-04-04 M. AbuGhanem

Digital Signal Processing functions are widely used in real time high speed applications. Those functions are generally implemented either on ASICs with inflexibility, or on FPGAs with bottlenecks of relatively smaller utilization factor or…

Other Computer Science · Computer Science 2013-06-04 Amitabha Sinha , Soumojit Acharyya , Suranjan Chakraborty , Mitrava Sarkar

Neural network accelerators have been widely applied to edge devices for complex tasks like object tracking, image recognition, etc. Previous works have explored the quantization technologies in related lightweight accelerator designs to…

Hardware Architecture · Computer Science 2026-02-27 Yuhao Liu , Salim Ullah , Akash Kumar

Compensating for nonlinear effects using digital signal processing (DSP) is complex and computationally expensive in long-haul optical communication systems due to intractable interactions between Kerr nonlinearity, chromatic dispersion…

Signal Processing · Electrical Eng. & Systems 2023-08-24 Naveenta Gautam , Sai Vikranth Pendem , Brejesh Lall , Amol Choudhary

The RFX-mod2 Nuclear Fusion experiment is an upgrade of RFX-mod, shutdown in 2016. Among the other improvements in the machine structure and diagnostics, a larger number of electromagnetic probes (EMs) is foreseen to provide more…

We present a one-step scheme to construct the controlled-phase gate deterministically on remote transmon qutrits coupled to different resonators connected by a superconducting transmission line for an universal distributed quantum…

Quantum Physics · Physics 2018-09-05 Ming Hua , Ming-Jie Tao , Ahmed Alsaedi , Tasawar Hayat , Fu-Guo Deng

This paper describes how we successfully used the HPX programming model to port the DCA++ application on multiple architectures that include POWER9, x86, ARM v8, and NVIDIA GPUs. We describe the lessons we can learn from this experience as…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-10-21 Weile Wei , Arghya Chatterjee , Kevin Huck , Oscar Hernandez , Hartmut Kaiser

FPGAs have been shown to be a promising platform for deploying Quantised Neural Networks (QNNs) with high-speed, low-latency, and energy-efficient inference. However, the complexity of modern deep-learning models limits the performance on…

Hardware Architecture · Computer Science 2025-11-06 Changhong Li , Biswajit Basu , Shreejith Shanker

For the first time in history, we are seeing a branching point in computing paradigms with the emergence of quantum processing units (QPUs). Extracting the full potential of computation and realizing quantum algorithms with a…

Quantum Physics · Physics 2022-11-29 Sergey Bravyi , Oliver Dial , Jay M. Gambetta , Dario Gil , Zaira Nazario

We present a shared memory implementation of a parallel algorithm, called delta-stepping, for solving the single source shortest path problem for directed and undirected graphs. In order to reduce synchronization costs we make some…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-02-21 M. Kranjčević , D. Palossi , S. Pintarelli

Adaptive Computation (AC) has been shown to be effective in improving the efficiency of Open-Domain Question Answering (ODQA) systems. However, current AC approaches require tuning of all model parameters, and training state-of-the-art ODQA…

Computation and Language · Computer Science 2021-07-06 Yuxiang Wu , Pasquale Minervini , Pontus Stenetorp , Sebastian Riedel

Here we investigate analogy between quantum signal processing (QSP) and the adiabatic-impulse model (AIM) in order to implement the QSP algorithm with fast quantum logic gates. QSP is an algorithm that uses single-qubit dynamics to perform…

Quantum Physics · Physics 2025-12-02 D. O. Shendryk , O. V. Ivakhnenko , S. N. Shevchenko , Franco Nori

Artificial intelligence (AI) has become a pivotal force in reshaping next generation mobile networks. Edge computing holds promise in enabling AI as a service (AIaaS) for prompt decision-making by offloading deep neural network (DNN)…

Networking and Internet Architecture · Computer Science 2025-01-28 Vahid Pourakbar , Hamed Shah-Mansouri

In this paper, we present two symbiotic optimizations to optimize recursive task parallel (RTP) programs by reducing the task creation and termination overheads. Our first optimization Aggressive Finish-Elimination (AFE) helps reduce the…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-02-24 Suyash Gupta , Rahul Shrivastava , V. Krishna Nandivada

We present a GPU-portable implementation of a real-space density functional theory (DFT) code ``QUMASUN'' and benchmark it on the new Plasma Simulator featuring Intel Xeon 6980P CPUs, and AMD MI300A GPUs. Additional tests were performed on…

Computational Physics · Physics 2025-12-08 Atsushi M. Ito

The ALICE experiment has undergone a major upgrade for LHC Run 3 and will collect data at an interaction rate 50 times larger than before. The new computing scheme for Run 3 replaces the traditionally separate online and offline frameworks…

Instrumentation and Detectors · Physics 2022-08-17 David Rohr

This paper proposes a new protocol called Optimal DCF (O-DCF). Inspired by a sequence of analytic results, O-DCF modifies the rule of adapting CSMA parameters, such as backoff time and transmission length, based on a function of the…

Networking and Internet Architecture · Computer Science 2012-07-17 Jinsung Lee , Yung Yi , Song Chong , Bruno Nardelli , Edward W. Knightly , Mung Chiang