English
Related papers

Related papers: Learning Quantized Continuous Controllers for Inte…

200 papers

Deep neural network (DNN)-based policy models like vision-language-action (VLA) models are transformative in automating complex decision-making across applications by interpreting multi-modal data. However, scaling these models greatly…

Robotics · Computer Science 2024-12-03 Seongmin Park , Hyungmin Kim , Wonseok Jeon , Juyoung Yang , Byeongwook Jeon , Yoonseon Oh , Jungwook Choi

Large speech recognition models like Whisper-small achieve high accuracy but are difficult to deploy on edge devices due to their high computational demand. To this end, we present a unified, cross-library evaluation of post-training…

Audio and Speech Processing · Electrical Eng. & Systems 2026-05-22 Arthur Söhler , Julian Irigoyen , Andreas Søeborg Kirkedal

Neural network quantization is a promising compression technique to reduce memory footprint and save energy consumption, potentially leading to real-time inference. However, there is a performance gap between quantized and full-precision…

Computer Vision and Pattern Recognition · Computer Science 2022-02-11 Qing Jin , Jian Ren , Richard Zhuang , Sumant Hanumante , Zhengang Li , Zhiyu Chen , Yanzhi Wang , Kaiyuan Yang , Sergey Tulyakov

State-space models (SSMs) have recently gained attention in deep learning for their ability to efficiently model long-range dependencies, making them promising candidates for edge-AI applications. In this paper, we analyze the effects of…

Machine Learning · Computer Science 2025-06-17 Leo Zhao , Tristan Torchet , Melika Payvand , Laura Kriener , Filippo Moro

In this paper, we propose MCUBERT to enable language models like BERT on tiny microcontroller units (MCUs) through network and scheduling co-optimization. We observe the embedding table contributes to the major storage bottleneck for tiny…

Machine Learning · Computer Science 2024-10-24 Zebin Yang , Renze Chen , Taiqiang Wu , Ngai Wong , Yun Liang , Runsheng Wang , Ru Huang , Meng Li

Post-Quantum Cryptographic (PQC) algorithms are mathematically secure and resistant to quantum attacks but can still leak sensitive information in hardware implementations due to natural faults or intentional fault injections. The intent…

Cryptography and Security · Computer Science 2025-08-06 Rourab Paul , Paresh Baidya , Krishnendu Guha

In this paper, we present a deep reinforcement learning platform named FIXAR which employs fixed-point data types and arithmetic units for the first time using a SW/HW co-design approach. Starting from 32-bit fixed-point data,…

Hardware Architecture · Computer Science 2021-02-25 Je Yang , Seongmin Hong , Joo-Young Kim

Quantized neural network (NN) with a reduced bit precision is an effective solution to reduces the computational and memory resource requirements and plays a vital role in machine learning. However, it is still challenging to avoid the…

Machine Learning · Computer Science 2020-10-23 Xiaobin Li , Hongxu Jiang , Shuangxi Huang , Fangzheng Tian

This paper presents an optimized methodology to design and deploy Speech Enhancement (SE) algorithms based on Recurrent Neural Networks (RNNs) on a state-of-the-art MicroController Unit (MCU), with 1+8 general-purpose RISC-V cores. To…

Sound · Computer Science 2022-10-17 Manuele Rusci , Marco Fariselli , Martin Croome , Francesco Paci , Eric Flamand

Convolutional Neural Networks (CNNs) have become common in many fields including computer vision, speech recognition, and natural language processing. Although CNN hardware accelerators are already included as part of many SoC…

Quantized low-precision neural networks are very popular because they require less computational resources for inference and can provide high performance, which is vital for real-time and embedded recognition systems. However, their…

Computer Vision and Pattern Recognition · Computer Science 2020-10-21 Anton Trusov , Elena Limonova , Dmitry Slugin , Dmitry Nikolaev , Vladimir V. Arlazarov

Quantized neural networks are well known for reducing the latency, power consumption, and model size without significant harm to the performance. This makes them highly appropriate for systems with limited resources and low power capacity.…

Machine Learning · Computer Science 2024-06-11 Moshe Kimhi , Tal Rozen , Avi Mendelson , Chaim Baskin

Multi-bit quantization networks enable flexible deployment of deep neural networks by supporting multiple precision levels within a single model. However, existing approaches suffer from significant training overhead as full-dataset updates…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Jinhee Kim , Jae Jun An , Kang Eun Jeon , Jong Hwan Ko

Pre-trained vision transformers have achieved remarkable performance across various visual tasks but suffer from expensive computational and memory costs. While model quantization reduces memory usage by lowering precision, these models…

Machine Learning · Computer Science 2025-08-06 Ching-Yi Lin , Sahil Shah

Over the past decade, machine learning techniques have revolutionized how research is done, from designing new materials and predicting their properties to assisting drug discovery to advancing cybersecurity. Recently, we added to this list…

Quantum Physics · Physics 2018-10-19 Justyna P. Zwolak , Sandesh S. Kalantre , Xingyao Wu , Stephen Ragole , Jacob M. Taylor

In order to solve problems of practical importance, quantum computers will likely need to incorporate quantum error correction, where a logical qubit is redundantly encoded in many noisy physical qubits. The large physical-qubit overhead…

Quantum Physics · Physics 2025-03-25 Harald Putterman , Kyungjoo Noh , Connor T. Hann , Gregory S. MacCabe , Shahriar Aghaeimeibodi , Rishi N. Patel , Menyoung Lee , William M. Jones , Hesam Moradinejad , Roberto Rodriguez , Neha Mahuli , Jefferson Rose , John Clai Owens , Harry Levine , Emma Rosenfeld , Philip Reinhold , Lorenzo Moncelsi , Joshua Ari Alcid , Nasser Alidoust , Patricio Arrangoiz-Arriola , James Barnett , Przemyslaw Bienias , Hugh A. Carson , Cliff Chen , Li Chen , Harutiun Chinkezian , Eric M. Chisholm , Ming-Han Chou , Aashish Clerk , Andrew Clifford , R. Cosmic , Ana Valdes Curiel , Erik Davis , Laura DeLorenzo , J. Mitchell D'Ewart , Art Diky , Nathan D'Souza , Philipp T. Dumitrescu , Shmuel Eisenmann , Essam Elkhouly , Glen Evenbly , Michael T. Fang , Yawen Fang , Matthew J. Fling , Warren Fon , Gabriel Garcia , Alexey V. Gorshkov , Julia A. Grant , Mason J. Gray , Sebastian Grimberg , Arne L. Grimsmo , Arbel Haim , Justin Hand , Yuan He , Mike Hernandez , David Hover , Jimmy S. C. Hung , Matthew Hunt , Joe Iverson , Ignace Jarrige , Jean-Christophe Jaskula , Liang Jiang , Mahmoud Kalaee , Rassul Karabalin , Peter J. Karalekas , Andrew J. Keller , Amirhossein Khalajhedayati , Aleksander Kubica , Hanho Lee , Catherine Leroux , Simon Lieu , Victor Ly , Keven Villegas Madrigal , Guillaume Marcaud , Gavin McCabe , Cody Miles , Ashley Milsted , Joaquin Minguzzi , Anurag Mishra , Biswaroop Mukherjee , Mahdi Naghiloo , Eric Oblepias , Gerson Ortuno , Jason Pagdilao , Nicola Pancotti , Ashley Panduro , JP Paquette , Minje Park , Gregory A. Peairs , David Perello , Eric C. Peterson , Sophia Ponte , John Preskill , Johnson Qiao , Gil Refael , Rachel Resnick , Alex Retzker , Omar A. Reyna , Marc Runyan , Colm A. Ryan , Abdulrahman Sahmoud , Ernesto Sanchez , Rohan Sanil , Krishanu Sankar , Yuki Sato , Thomas Scaffidi , Salome Siavoshi , Prasahnt Sivarajah , Trenton Skogland , Chun-Ju Su , Loren J. Swenson , Stephanie M. Teo , Astrid Tomada , Giacomo Torlai , E. Alex Wollack , Yufeng Ye , Jessica A. Zerrudo , Kailing Zhang , Fernando G. S. L. Brandão , Matthew H. Matheny , Oskar Painter

This study examines 4-bit quantization methods like GPTQ in large language models (LLMs), highlighting GPTQ's overfitting and limited enhancement in Zero-Shot tasks. While prior works merely focusing on zero-shot measurement, we extend task…

Ultra-low-precision inference can sharply reduce memory and latency but often degrades accuracy and relies on specialized hardware. We present SONIQ, a system-optimized, noise-injected quantization framework that learns per-channel mixed…

Hardware Architecture · Computer Science 2025-11-11 Cyrus Zhou , Pedro Savarese , Zack Hassman , Vaughn Richard , Michael DiBrino , Michael Maire , Yanjing Li

As quantum hardware advances toward fault-tolerant operation, an intermediate stage known as early fault-tolerant quantum computing (EFTQC) is emerging, where partial error correction enables meaningful computation. In this regime, the…

Quantum Physics · Physics 2025-11-14 Yanbing Zhou , Athena Caesura , Corneliu Buda , Xavier Jackson , Clena M. Abuan , Shangjie Guo

The massive computational costs associated with large language model (LLM) pretraining have spurred great interest in reduced-precision floating-point representations to accelerate the process. As a result, the BrainFloat16 (BF16) precision…

Machine Learning · Computer Science 2025-03-26 Joonhyung Lee , Jeongin Bae , Byeongwook Kim , Se Jung Kwon , Dongsoo Lee