Related papers: Automatic Hybrid-Precision Quantization for MIMO D…
Effective access points (APs) selection is a crucial step in localization systems. It directly affects both localization accuracy and computational efficiency. Classical APs selection algorithms are usually computationally expensive,…
We address the problem of network quantization, that is, reducing bit-widths of weights and/or activations to lighten network architectures. Quantization methods use a rounding function to map full-precision values to the nearest quantized…
Multiple-input multiple-output (MIMO) is critical for 6G communication, offering improved spectral efficiency and reliability. However, conventional fully digital designs face significant challenges due to high hardware complexity and power…
As quantum networks evolve toward a full quantum Internet, reliable transmission in quantum multiple-input multiple-output (QuMIMO) settings becomes essential, yet remains difficult due to noise, crosstalk, and the mixing of quantum…
Weight quantization effectively reduces memory consumption and enable the deployment of Large Language Models on edge devices, yet existing hardware-friendly methods often rely on uniform quantization, which suffers from poor…
The growing scale of large language models (LLMs) not only demands extensive computational resources but also raises environmental concerns due to their increasing carbon footprint. Model quantization emerges as an effective approach that…
Having lower quantization resolution, has been introduced in the literature, as a solution to reduce the power consumption of massive MIMO and millimeter wave MIMO systems. In this paper, we analyze bit error rate (BER) performance of…
Mixed-precision quantization (MPQ) suffers from the time-consuming process of searching the optimal bit-width allocation i.e., the policy) for each layer, especially when using large-scale datasets such as ISLVRC-2012. This limits the…
We describe a measure quantization procedure i.e., an algorithm which finds the best approximation of a target probability law (and more generally signed finite variation measure) by a sum of $Q$ Dirac masses ($Q$ being the quantization…
Large speech recognition models like Whisper-small achieve high accuracy but are difficult to deploy on edge devices due to their high computational demand. To this end, we present a unified, cross-library evaluation of post-training…
Mixed-Precision Quantization~(MQ) can achieve a competitive accuracy-complexity trade-off for models. Conventional training-based search methods require time-consuming candidate training to search optimized per-layer bit-width…
Large Language Models (LLMs) have revolutionized natural language processing tasks. However, their practical application is constrained by substantial memory and computational demands. Post-training quantization (PTQ) is considered an…
Deploying Large Language Models (LLMs) on resource-constrained edge devices like the Raspberry Pi presents challenges in computational efficiency, power consumption, and response latency. This paper explores quantization-based optimization…
We consider the weak target detection problem with unknown parameter in colocated multiple-input multiple-output (MIMO) radar. To cope with the sheer amount of data for large-size systems, a multi-bit quantizer is utilized in the sampling…
Dynamic runtime latency and memory constraints necessitate flexible large language model (LLM) deployment, where an LLM can be inferred with various quantization precisions based on available computational resources. Recent work on such…
Quantizing deep convolutional neural networks for image super-resolution substantially reduces their computational costs. However, existing works either suffer from a severe performance drop in ultra-low precision of 4 or lower bit-widths,…
The sub-THz spectrum offers numerous advantages, including massive multiple-input multiple-output (MIMO) technology with large antenna arrays that enhance spectral efficiency (SE) of future systems. Hybrid precoding (HP) thus emerges as a…
The high hardware complexity of a massive MIMO base station, which requires hundreds of radio chains, makes it challenging to build commercially. One way to reduce the hardware complexity and power consumption of the receiver is to lower…
The ever-growing computational complexity of Large Language Models (LLMs) necessitates efficient deployment strategies. The current state-of-the-art approaches for Post-training Quantization (PTQ) often require calibration to achieve the…
As the number of IoT devices continue to exponentially increase and saturate the wireless spectrum, there is a dire need for additional spectrum to support large networks of wireless devices. Over the past years, many promising solutions…