English
Related papers

Related papers: Wilson and Domainwall Kernels on Oakforest-PACS

200 papers

Increasing data center network speed coupled with application requirements for high throughput and low latencies have raised the efficiency bar for network stacks. To reduce substantial kernel overhead in network processing, recent…

Operating Systems · Computer Science 2023-03-02 Yihan Yang , Zhuobin Huang , Antoine Kaufmann , Jialin Li

The Crossroads supercomputer was designed to simulate some of the most complex physical devices in the world. These simulations routinely require 1/2 petabyte or more of system memory running on thousands of compute nodes for months at a…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-11-11 Galen M. Shipman , Sriram Swaminarayan , Gary Grider , Jim Lujan , R. Joseph Zerr

Recently the use of neural networks has been introduced in the context of the signed particle formulation of quantum mechanics to rapidly and reliably compute the Wigner kernel of any provided potential. This new technique has introduced…

Computational Physics · Physics 2018-06-04 Jean Michel Sellier , Jacob Leygonie , Gaetan Marceau Caron

For data analysis of large-scale experiments such as LHC Atlas and other Japanese high energy and nuclear physics projects, we have constructed a Grid test bed at ICEPP and KEK. These institutes are connected to national scientific gigabit…

We propose a new protocol to implement ultra-fast two-qubit phase gates with trapped ions using spin-dependent kicks induced by resonant transitions. By only optimizing the allocation of the arrival times in a pulse train sequence the gate…

Quantum Physics · Physics 2020-10-28 E. Torrontegui , D. Heinrich , M. I. Hussain , R. Blatt , J. J. García-Ripoll

The well-known Smith-Waterman (SW) algorithm is the most commonly used method for local sequence alignments. However, SW is very computationally demanding for large protein databases. There exist several implementations that take advantage…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-02-24 Enzo Rucci , Carlos Garcia , Guillermo Botella , Armando De Giusti , Marcelo Naiouf , Manuel Prieto-Matias

Existing high-performance computing (HPC) interconnection architectures are based on high-radix switches, which limits the injection/local performance and introduces latency/energy/cost overhead. The new wafer-scale packaging and high-speed…

Hardware Architecture · Computer Science 2024-08-27 Yinxiao Feng , Kaisheng Ma

With recent developments in parallel supercomputing architecture, many core, multi-core, and GPU processors are now commonplace, resulting in more levels of parallelism, memory hierarchy, and programming complexity. It has been necessary to…

High Energy Physics - Lattice · Physics 2017-12-04 Ruizi Li , Carleton DeTar , Steven Gottlieb , Doug Toussaint

Deep Neural Networks are becoming increasingly popular in always-on IoT edge devices performing data analytics right at the source, reducing latency as well as energy consumption for data communication. This paper presents CMSIS-NN,…

Neural and Evolutionary Computing · Computer Science 2018-01-23 Liangzhen Lai , Naveen Suda , Vikas Chandra

Modern OpenMP threading techniques are used to convert the MPI-only Hartree-Fock code in the GAMESS program to a hybrid MPI/OpenMP algorithm. Two separate implementations that differ by the sharing or replication of key data structures…

Distributed, Parallel, and Cluster Computing · Computer Science 2017-08-15 Vladimir Mironov , Yuri Alexeev , Kristopher Keipert , Michael D'mello , Alexander Moskovsky , Mark S. Gordon

With the ease-of-programming, flexibility and yet efficiency, MapReduce has become one of the most popular frameworks for building big-data applications. MapReduce was originally designed for distributed-computing, and has been extended to…

Distributed, Parallel, and Cluster Computing · Computer Science 2013-09-03 Mian Lu , Lei Zhang , Huynh Phung Huynh , Zhongliang Ong , Yun Liang , Bingsheng He , Rick Siow Mong Goh , Richard Huynh

Lightweight Tunnels (LWTs) in the Linux kernel enable efficient per-route tunneling and are widely used by protocols such as In Situ Operations, Administration, and Maintenance (IOAM), Segment Routing over IPv6 (SRv6), and Routing Protocol…

Networking and Internet Architecture · Computer Science 2025-03-20 J. Iurman , E. Wansart , M. Goffart , B. Donnet

The most computationally demanding part of Lattice QCD simulations is solving quark propagators. Quark propagators are typically obtained with a linear equation solver utilizing HPC machines. The CCS QCD Benchmark is a benchmark program…

High Energy Physics - Lattice · Physics 2017-01-01 Taisuke Boku , Ken-Ichi Ishikawa , Yoshinobu Kuramashi , Lawrence Meadows , Michael D`Mello , Maurice Troute , Ravi Vemuri

The introduction of Intel(R) Xeon Phi(TM) coprocessors opened up new possibilities in development of highly parallel applications. The familiarity and flexibility of the architecture together with compiler support integrated into the Intel…

Distributed, Parallel, and Cluster Computing · Computer Science 2012-11-26 Jiri Dokulil , Enes Bajrovic , Siegfried Benkner , Sabri Pllana , Martin Sandrieser , Beverly Bachmayer

In addition to hardware wall-time restrictions commonly seen in high-performance computing systems, it is likely that future systems will also be constrained by energy budgets. In the present work, finite difference algorithms of varying…

Mathematical Software · Computer Science 2017-09-29 Satya P. Jammy , Christian T. Jacobs , David J. Lusher , Neil D. Sandham

In this technical report, we introduce FastFusionNet, an efficient variant of FusionNet [12]. FusionNet is a high performing reading comprehension architecture, which was designed primarily for maximum retrieval accuracy with less regard…

Computation and Language · Computer Science 2019-03-05 Felix Wu , Boyi Li , Lequn Wang , Ni Lao , John Blitzer , Kilian Q. Weinberger

In quantum kernel learning, the primary method involves using a quantum computer to calculate the inner product between feature vectors, thereby obtaining a Gram matrix used as a kernel in machine learning models such as support vector…

Quantum Physics · Physics 2024-05-17 Hiroshi Yamauchi , Tomah Sogabe , Rodney Van Meter

This paper reports our efforts on swCaffe, a highly efficient parallel framework for accelerating deep neural networks (DNNs) training on Sunway TaihuLight, the current fastest supercomputer in the world that adopts a unique many-core…

Distributed, Parallel, and Cluster Computing · Computer Science 2019-03-19 Jiarui Fang , Liandeng Li , Haohuan Fu , Jinlei Jiang , Wenlai Zhao , Conghui He , Xin You , Guangwen Yang

Data centers based on Passive Optical Networks (PONs) can offer scalability, low cost and high energy-efficiency. Application in data centers can use Virtual Machines (VMs) to provide efficient utilization of the physical resources. This…

Networking and Internet Architecture · Computer Science 2022-03-25 Mohammed Alharthi , Sanaa H. Mohamed , Barzan Yosuf , Taisir E. H. El-Gorashi , Jaafar M. H. Elmirghani

Using WiFi signals for indoor localization is the main localization modality of the existing personal indoor localization systems operating on mobile devices. WiFi fingerprinting is also used for mobile robots, as WiFi signals are usually…

Robotics · Computer Science 2017-05-01 Michał Nowicki , Jan Wietrzykowski