Related papers: Optimizing the domain wall fermion Dirac operator …
Artificial intelligence necessitates adaptable hardware accelerators for efficient high-throughput million operations. We present pipelined architecture with CORDIC block for linear MAC computations and nonlinear iterative Activation…
CFD users of supercomputers usually resort to rule-of-thumb methods to select the number of subdomains (partitions) when relying on MPI-based parallelization. One common approach is to set a minimum number of elements or cells per…
Domain-specific systems-on-chip (DSSoCs) aim at bridging the gap between application-specific integrated circuits (ASICs) and general-purpose processors. Traditional operating system (OS) schedulers can undermine the potential of DSSoCs…
COSA is a novel CFD system based on the compressible Navier-Stokes model for unsteady aerodynamics and aeroelasticity of fixed structures, rotary wings and turbomachinery blades. It includes a steady, time domain, and harmonic balance flow…
Fault injection attacks deliberately inject faults into a device via physical channels to disturb its regular execution. Adversaries can effectively deduce secrets by analyzing both the normal and faulty outputs, posing serious threats to…
This paper describes a state-of-the-art parallel Lattice QCD Monte Carlo code for staggered fermions, purposely designed to be portable across different computer architectures, including GPUs and commodity CPUs. Portability is achieved…
We outline a strategy to compute deeply inelastic scattering structure functions using a hybrid quantum computer. Our approach takes advantage of the representation of the fermion determinant in the QCD path integral as a quantum mechanical…
Configuring stream processing systems for efficient performance, especially in cloud-native deployments, is a challenging and largely manual task. We present an experiment-driven approach for automated configuration optimization that…
We propose a novel, machine-learning-based framework for constructing lattice fermions using Physics-Informed Neural Networks (PINNs). Our approach treats the formulation of the Dirac operator as an optimization problem guided by physical…
Load balancing, operator instance collocations and horizontal scaling are critical issues in Parallel Stream Processing Engines to achieve low data processing latency, optimized cluster utilization and minimized communication cost…
In this paper, we propose a cross subcarrier precoder design (CSPD) for massive multiple-input multiple-output (MIMO) orthogonal frequency division multiplexing (OFDM) systems. The aim is to maximize the weighted sum-rate (WSR) performance…
We present our recent studies of the pseudo-critical temperature, $T_c$, of QCD using domain wall fermions. Domain wall fermions have the advantage that they preserve exact SU(2) chiral symmetry at finite lattice spacing in the limit that…
High throughput and low latency data processing is essential for systems requiring live decision making, control, and machine learning-optimized data reduction. We focus on two distinct use cases for in-flight streaming data processing for…
We determine conditions for the quantisation of graphs using the Dirac operator for both two and four component spinors. According to the Bohigas-Giannoni-Schmit conjecture for such systems with time-reversal symmetry the energy level…
Block operations during simulation in Scicos and VSS environments can naturally be described as Nsp functions. But the direct use of Nsp functions for simulation leads to poor performance since the Nsp language is interpreted, not compiled.…
Ferromagnetic resonance is used to reveal features of the buried electronic band structure at interfaces between ferromagnetic metals and topological insulators. By monitoring the evolution of magnetic damping, the application of this…
We present a method for optimizing the location of the fermion ground-state nodes using a combination of diffusion Monte Carlo (DMC) and projected gradient descent (PGD). A PGD iteration shifts the parameters of an arbitrary node-fixing…
With the increasing digital services demand, performance and power-efficiency become vital requirements for digital circuits and systems. However, the enabling CMOS technology scaling has been facing significant challenges of device…
To tackle combinatorial optimization problems using an Ising machine, the objective function and constraints must be mapped onto a quadratic unconstrained binary optimization (QUBO) model. While QUBO involves binary variables, combinatorial…
Particle Swarm Optimization technique offers optimal or suboptimal solution to multidimensional rough objective functions. In this paper, this optimization technique is used for designing fractional order PID controllers that give better…