Related papers: Real-Time Data Processing in the Muon System of th…
To collect data for the study of charm particle decays, we built a high speed data acquisition system for use with the E791 magnetic spectrometer at Fermilab. The DA system read out 24000 channels in 50 uS. Events were accepted at the rate…
Deploying mixed-precision neural networks on edge devices is friendly to hardware resources and power consumption. To support fully mixed-precision neural network inference, it is necessary to design flexible hardware accelerators for…
A high-intensity pulsed muon beam is becoming available at the at the Japan Proton Accelerator Research Complex (J-PARC). Many experiments to study fundamental physics using this high-intensity muon beam are proposed. An experiment to…
The escalating data volume and complexity resulting from the rapid expansion of artificial intelligence (AI), internet of things (IoT) and 5G/6G mobile networks is creating an urgent need for energy-efficient, scalable computing hardware.…
Single muon triggers are crucial for the physics programmes at hadron collider experiments. To be sensitive to electroweak processes, single muon triggers with transverse momentum thresholds down to 20 GeV and dimuon triggers with even…
In recent years, transformer models have revolutionized Natural Language Processing (NLP) and shown promising performance on Computer Vision (CV) tasks. Despite their effectiveness, transformers' attention operations are hard to accelerate…
The performances of automatic speech recognition (ASR) systems degrade drastically under noisy conditions. Explicit distortion modelling (EDM), as a feature compensation step, is able to enhance ASR systems under such conditions by…
We present the results of a search for the effects of large extra spatial dimensions in $p{\bar p}$ collisions at $\sqrt{s} =$ 1.96 TeV in events containing a pair of energetic muons. The data correspond to 246 \ipb of integrated luminosity…
Attention-based large language models (LLMs) have transformed modern AI applications, but the quadratic cost of self-attention imposes significant compute and memory overhead. Dynamic sparsity (DS) attention mitigates this, yet its hardware…
The heaviest known Fermion particle -- the top quark -- was discovered at Fermilab in the first run of the Tevatron in 1995. However, besides its mere existence one needs to study its properties precisely in order to verify or falsify the…
Memory bandwidth has become the real-time bottleneck of current deep learning accelerators (DLA), particularly for high definition (HD) object detection. Under resource constraints, this paper proposes a low memory traffic DLA chip with…
We present the upgraded design, construction, and beam test results for the Muon Trigger Detector (MTD) developed for the muon Electric Dipole Moment (muEDM) experiment at the Paul Scherrer Institute (PSI) in Switzerland. This experiment…
The Muon to Electron Experiment (Mu2e) requires a uniform beam profile from the Muon Delivery Ring to meet their experimental needs. A specialized Spill Regulation System (SRS) has been developed to help achieve consistent spill uniformity.…
Digital Signal Processing functions are widely used in real time high speed applications. Those functions are generally implemented either on ASICs with inflexibility, or on FPGAs with bottlenecks of relatively smaller utilization factor or…
The recent surge of interest in Deep Neural Networks (DNNs) has led to increasingly complex networks that tax computational and memory resources. Many DNNs presently use 16-bit or 32-bit floating point operations. Significant performance…
In this article we study the suitability of dierent computational accelerators for the task of real-time data processing. The algorithm used for comparison is the polyphase filter, a standard tool in signal processing and a well established…
The Muon g-2 experiment at Fermilab, with the aim to measure the muon anomalous magnetic moment to an unprecedented level of 140~ppb, has started beam and detector commissioning in Summer 2017. To deal with incoming data projected to be…
The Fermilab Tevatron (p pbar), operating at sqrt(s)=1.96 TeV, is a rich source of B hadrons. The large acceptance in terms of rapidity and transverse momentum of the charged particle tracking system and the muon system make the upgraded…
HIPSR (HI-Pulsar) is a digital signal processing system for the Parkes 21-cm Multibeam Receiver that provides larger instantaneous bandwidth, increased dynamic range, and more signal processing power than the previous systems in use at…
A multiply-accumulate (MAC) operation is the main computation unit for DSP applications. DSP blocks are one of the efficient solutions to implement MACs in FPGA's. However, since the DSP blocks have wide multiplier and adder blocks, MAC…