English
Related papers

Related papers: Wilson and Domainwall Kernels on Oakforest-PACS

200 papers

I describe here the performance of a parallel treecode with individual particle timesteps. The code is based on the Barnes-Hut algorithm and runs cosmological N-body simulations on parallel machines with a distributed memory architecture…

Astrophysics · Physics 2009-11-07 R. Valdarnini

Hardware technological advances are struggling to match scientific ambition, and a key question is how we can use the transistors that we already have more effectively. This is especially true for HPC, where the tendency is often to throw…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-11-11 Nick Brown

The deployment of Quantized Neural Networks (QNN) on advanced microcontrollers requires optimized software to exploit digital signal processing (DSP) extensions of modern instruction set architectures (ISA). As such, recent research…

Hardware Architecture · Computer Science 2020-07-16 Nazareno Bruschi , Angelo Garofalo , Francesco Conti , Giuseppe Tagliavini , Davide Rossi

The rapid expansion of Internet of Things (IoT) deployments has enlarged the attack surface of modern digital infrastructure while exposing a key security mismatch: many intrusion detection systems (IDSs) remain too computationally…

Cryptography and Security · Computer Science 2026-05-06 Dileepa Mabulage , Banuka Athuraliya

To tackle combinatorial optimization problems using an Ising machine, the objective function and constraints must be mapped onto a quadratic unconstrained binary optimization (QUBO) model. While QUBO involves binary variables, combinatorial…

Statistical Mechanics · Physics 2025-01-22 Shuta Kikuchi , Kotaro Takahashi , Shu Tanaka

We present spintronic devices based hardware implementation of UNet for segmentation tasks. Our approach involves designing hardware for convolution, deconvolution, rectified activation function (ReLU), and max pooling layers of the UNet…

Emerging Technologies · Computer Science 2024-07-12 Venkatesh Vadde , Bhaskaran Muralidharan , Abhishek Sharma

The wireless network places vital role in the present day communication scenario. The ad hoc nature of wireless communication adds flavour to suit various real world applications. This improves the performance of the network tremendously…

Networking and Internet Architecture · Computer Science 2013-01-15 S. Thirumurugan , E. George Dharma Prakash Raj

Programming efficiently heterogeneous systems is a major challenge, due to the complexity of their architectures. Intel oneAPI, a new and powerful standards-based unified programming model, built on top of SYCL, addresses these issues. In…

Distributed, Parallel, and Cluster Computing · Computer Science 2021-09-16 Raúl Nozal , Jose Luis Bosque

Supervised learning of Convolutional Neural Networks (CNNs), also known as supervised Deep Learning, is a computationally demanding process. To find the most suitable parameters of a network for a given application, numerous training…

Distributed, Parallel, and Cluster Computing · Computer Science 2015-07-01 Andre Viebke , Sabri Pllana

Physics-informed neural networks (PINNs) are a class of deep learning models that utilize physics in the form of differential equations to address complex problems, including those with limited data availability. However, solving…

Machine Learning · Computer Science 2026-03-26 Himanshu Pandey , Anshima Singh , Ratikanta Behera

Quantum computing has garnered attention for its potential to solve complex computational problems with considerable speedup. Despite notable advancements in the field, achieving meaningful scalability and noise control in quantum hardware…

Quantum Physics · Physics 2025-05-12 Eduardo Willwock Lussi , Rafael de Santiago , Eduardo Inacio Duzzioni

Point cloud registration serves as a basis for vision and robotic applications including 3D reconstruction and mapping. Despite significant improvements on the quality of results, recent deep learning approaches are computationally…

Robotics · Computer Science 2024-04-02 Keisuke Sugiura , Hiroki Matsutani

Accelerated computing is widely used in high-performance computing. Therefore, it is crucial to experiment and discover how to better utilize GPUGPUs latest generations on relevant applications. In this paper, we present results and share…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-08-13 Baodi Shan , Mauricio Araya-Polo

Neuromorphic computing (NC) architecture has shown its suitability for energy-efficient computation. Amongst several systems, spin-orbit torque (SOT) based domain wall (DW) devices are one of the most energy-efficient contenders for NC. To…

Mesoscale and Nanoscale Physics · Physics 2022-12-16 Durgesh Kumar , Ramu Maddu , Hong Jing Chung , Hasibur Rahaman , Tianli Jin , Sabpreet Bhatti , Sze Ter Lim , Rachid Sbiaa , S. N. Piramanayagam

In this work, we first characterize the hybrid execution patterns of GCNs on Intel Xeon CPU. Guided by the characterization, we design a GCN accelerator, HyGCN, using a hybrid architecture to efficiently perform GCNs. Specifically, first,…

Distributed, Parallel, and Cluster Computing · Computer Science 2020-01-09 Mingyu Yan , Lei Deng , Xing Hu , Ling Liang , Yujing Feng , Xiaochun Ye , Zhimin Zhang , Dongrui Fan , Yuan Xie

AI power demand is growing at an unprecedented rate while power grids are often ailing and struggle to keep up. Grid expansion comes with high capital expenditure and long-distance transmission losses, yet there is abundant renewable energy…

Co-expression network is a critical technique for the identification of inter-gene interactions, which usually relies on all-pairs correlation (or similar measure) computation between gene expression profiles across multiple samples.…

Distributed, Parallel, and Cluster Computing · Computer Science 2016-09-28 Yongchao Liu , Tony Pan , Srinivas Aluru

As hardware architectures are evolving in the push towards exascale, developing Computational Science and Engineering (CSE) applications depend on performance portable approaches for sustainable software development. This paper describes…

Unikernels are famous for providing excellent performance in terms of boot times, throughput and memory consumption, to name a few metrics. However, they are infamous for making it hard and extremely time consuming to extract such…

Deep neural networks have achieved impressive results in computer vision and machine learning. Unfortunately, state-of-the-art networks are extremely compute and memory intensive which makes them unsuitable for mW-devices such as IoT…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-03-15 Renzo Andri , Lukas Cavigelli , Davide Rossi , Luca Benini
‹ Prev 1 8 9 10 Next ›