中文
相关论文

相关论文: PolarQuant: Leveraging Polar Transformation for Ef…

200 篇论文

The polarization decomposition of arbitrary binary-input memoryless channels (BMCs) is studied in this work. By introducing the polarization factor (PF), defined in terms of the conditional entropy of the channel output under various input…

信息论 · 计算机科学 2025-05-06 Tianfu Qi , Jun Wang

Arikan's recursive code construction is designed to polarize a collection of memoryless channels into a set of good and a set of bad channels, and it can be efficiently decoded using successive cancellation. It was recently shown that the…

信息论 · 计算机科学 2019-01-16 Benjamin Bourassa , Maxime Tremblay , David Poulin

Quantum memory for flying optical qubits is a key enabler for a wide range of applications in quantum information science and technology. A critical figure of merit is the overall storage-and-retrieval efficiency. So far, despite the recent…

量子物理 · 物理学 2018-02-02 P. Vernaz-Gris , K. Huang , M. Cao , A. S. Sheremet , J. Laurat

Quantization can be used to form new vectors/matrices with shared values close to the original. In recent years, the popularity of scalar quantization for value-sharing applications has been soaring as it has been found huge utilities in…

机器学习 · 计算机科学 2019-12-11 Chen Wang , Xiaomei Yang , Shaomin Fei , Kai Zhou , Xiaofeng Gong , Miao Du , Ruisen Luo

In 2008 Arikan proposed polar coding [arXiv:0807.3917] which we summarize as follows: (a) From the root channel $W$ synthesize recursively a series of channels $W_N^{(1)},\dotsc,W_N^{(N)}$. (b) Select sophisticatedly a subset $A$ of…

信息论 · 计算机科学 2018-06-08 Hsin-Po Wang , Iwan Duursma

Large language models (LLMs) have shown strong performance across diverse tasks, but their inference with long input contexts is bottlenecked by memory size and bandwidth. The Key-Value (KV) cache size grows linearly with sequence length…

机器学习 · 计算机科学 2026-05-12 Junkai Zhang , Hang Guo , Luca Benini , Yawei Li

Large Language Models (LLMs) suffer inference-time memory bottlenecks dominated by the attention Key-Value (KV) cache, which scales with model size and context length. While KV-cache quantization alleviates this cost, bit allocation between…

机器学习 · 计算机科学 2026-05-12 Mohsen Hariri , Alan Luo , Weicong Chen , Shaochen Zhong , Tianyi Zhang , Qifan Wang , Xia Hu , Xiaotian Han , Vipin Chaudhary

Weight quantisation is an essential technique for enabling efficient training and deployment of modern deep learning models. However, the recipe book of quantisation formats is large and formats are often chosen empirically. In this paper,…

机器学习 · 计算机科学 2026-02-16 Douglas Orr , Luka Ribar , Carlo Luschi

Polar encoding, described by Arikan in IEEE Transactions on Information Theory, Vol. 55, No. 7, July 2009, was a milestone for telecommunications. A Polar code distributes information among high and low-capacity channels, showing the…

信息论 · 计算机科学 2025-07-29 Geraldo A. Barbosa

Polar coding is a recently proposed coding technique that can provably achieve the channel capacity. The polar code structure, which is based on the original 2x2 generator matrix, polarises the channels, i.e., a portion of the channel…

信息论 · 计算机科学 2019-01-08 Berksan Serbetci , Ali Emre Pusane

This paper describes efficient algorithms for computing rank-revealing factorizations of matrices that are too large to fit in RAM, and must instead be stored on slow external memory devices such as solid-state or spinning disk hard drives…

数学软件 · 计算机科学 2020-03-05 Nathan Heavner , Per-Gunnar Martinsson , Gregorio Quintana-Ortí

Video large language models (VideoLLMs) have demonstrated the capability to process longer video inputs and enable complex reasoning and analysis. However, due to the thousands of visual tokens from the video frames, the key-value (KV)…

计算机视觉与模式识别 · 计算机科学 2025-09-30 Keda Tao , Haoxuan You , Yang Sui , Can Qin , Huan Wang

Quantum Computing in the Noisy Intermediate-Scale Quantum (NISQ) era has shown promising applications in machine learning, optimization, and cryptography. Despite the progress, challenges persist due to system noise, errors, and decoherence…

量子物理 · 物理学 2024-04-12 Bikram Khanal , Pablo Rivas

Network quantization has emerged as one of the most practical model compression techniques, which significantly reduces a model's memory and compute consumption by mapping floating-point numbers to low-bit representations. However, existing…

计算机视觉与模式识别 · 计算机科学 2026-05-19 Peilin Sun , Jianxin Wu

Quantization of foundational models (FMs) is significantly more challenging than traditional DNNs due to the emergence of large magnitude values called outliers. Existing outlier-aware algorithm-architecture co-design techniques either use…

硬件体系结构 · 计算机科学 2025-05-01 Akshat Ramachandran , Souvik Kundu , Tushar Krishna

Applications of massive machine-type communications, such as sensor networks, smart metering, 'internet-of-things', or process and factory automation, are forecast to have great economic impact in the next five to ten years. Low-complexity…

信息论 · 计算机科学 2019-02-28 Joachim Neu

It has been shown that an extension of the basic binary polar transformation also polarizes over finite fields. With it the direct encoding of q-ary sources and channels is a process that can be implemented with simple and efficient…

信息论 · 计算机科学 2016-05-24 Ángel Bravo-Santos

Although LLM inference has emerged as a critical workload for many downstream applications, efficiently inferring LLMs is challenging due to the substantial memory footprint and bandwidth requirements. In parallel, compute capabilities have…

In real world, our datasets often contain outliers. Moreover, the outliers can seriously affect the final machine learning result. Most existing algorithms for handling outliers take high time complexities (e.g. quadratic or cubic…

计算几何 · 计算机科学 2020-02-28 Hu Ding , Zixiu Wang

Quantization emerges as one of the most promising compression technologies for deploying efficient large models for various real time application in recent years. Considering that the storage and IO of weights take up the vast majority of…

机器学习 · 计算机科学 2024-04-22 Yi Guo , Fanliu Kong , Xiaoyang Li , Hui Li , Wei Chen , Xiaogang Tian , Jinping Cai , Yang Zhang , Shouda Liu