Related papers: AmpliconHunter2: a SIMD-Accelerated In-Silico PCR …
This paper proposes a hardware-efficient architecture, Linearized Convolution Network (LiCo-Net) for keyword spotting. It is optimized specifically for low-power processor units like microcontrollers. ML operators exhibit heterogeneous…
Speech self-supervised models such as wav2vec 2.0 and HuBERT are making revolutionary progress in Automatic Speech Recognition (ASR). However, they have not been totally proven to produce better performance on tasks other than ASR. In this…
A search for dark matter axions with masses $>10 \mu eV/c^{2}$ has been performed using the HAYSTAC experiment's squeezed state receiver to achieve sub-quantum limited noise. This report includes details of the design and operation of the…
In LHC Run 3, ALICE will increase the data taking rate significantly to 50\,kHz continuous read out of minimum bias Pb-Pb events. This challenges the online and offline computing infrastructure, requiring to process 50 times as many events…
We introduce a state-of-the-art real-time, high-fidelity, audio codec leveraging neural networks. It consists in a streaming encoder-decoder architecture with quantized latent space trained in an end-to-end fashion. We simplify and speed-up…
A monolithic silicon pixel ASIC prototype, produced in 2024 as part of the Horizon 2020 MONOLITH ERC Advanced project, was tested with a 120 GeV/c pion beam. The ASIC features a matrix of hexagonal pixels with a 100 \mu m pitch, read by…
Recently, Transformer-based encoder-decoder models have demonstrated strong performance in multilingual speech recognition. However, the decoder's autoregressive nature and large size introduce significant bottlenecks during inference.…
Real-time text-driven joint audio-video avatar generation requires jointly synthesizing portrait video and speech with high fidelity and precise synchronization, yet existing audio-visual diffusion models remain too slow for interactive use…
We present the very first demonstration of a maser utilizing silicon vacancies (VSi) within 4H silicon carbide (SiC). Leveraging an innovative feedback-loop technique, we elevate the resonator's quality factor, enabling maser operation even…
This paper introduces wav2letter++, the fastest open-source deep learning speech recognition framework. wav2letter++ is written entirely in C++, and uses the ArrayFire tensor library for maximum efficiency. Here we explain the architecture…
Large language models (LLMs) have demonstrated remarkable abilities in natural language processing. However, their deployment on resource-constrained embedded devices remains difficult due to memory and computational demands. In this paper,…
We introduce a revised derivation of the bitwise Markov Chain Monte Carlo (MCMC) multiple-input multiple-output (MIMO) detector. The new approach resolves the previously reported high SNR stalling problem of MCMC without the need for…
This paper proposes a mechanism to accelerate and optimize the energy consumption of a face detection software based on Haar-like cascading classifiers, taking advantage of the features of low-cost Asymmetric Multicore Processors (AMPs)…
To search for dark matter candidates with masses below $\mathcal{O}$(MeV), the SPLENDOR (Search for Particles of Light dark mattEr with Narrow-gap semiconDuctORs) experiment is developing novel narrow-bandgap semiconductors with electronic…
At the Large Hadron Collider, the vast amount of data from experiments demands not only sophisticated algorithms but also substantial computational power for efficient processing. This paper introduces hardware acceleration as an essential…
We present a programmable 16 channel, mixed signal, low power readout ASIC, having the project historically named Gigasample Recorder of Analog waveforms from a PHotodetector (GRAPH). It is designed to read large aperture single photon…
Extremely low-bit quantization is critical for efficiently deploying Large Language Models (LLMs), yet it often leads to severe performance degradation at 2 bits and even at 4 bits (e.g., MXFP4). We present SignRoundV2, a post-training…
High Bandwidth Memory (HBM) provides massive aggregated memory bandwidth by exposing multiple memory channels to the processing units. To achieve high performance, an accelerator built on top of an FPGA configured with HBM (i.e., FPGA-HBM…
In this letter, we consider the uplink of a cell-free Massive multiple-input multiple-output (MIMO) network where each user is decoded by a subset of access points (APs). An additional step is introduced in the cell-free Massive MIMO…
The Picosecond Avalanche Detector is a multi-junction silicon pixel detector based on a $\mathrm{(NP)_{drift}(NP)_{gain}}$ structure, devised to enable charged-particle tracking with high spatial resolution and picosecond time-stamp…