English
Related papers

Related papers: BitHEP -- The Limits of Low-Precision ML in HEP

200 papers

Autoregressive decoding with generative Large Language Models (LLMs) on accelerators (GPUs/TPUs) is often memory-bound where most of the time is spent on transferring model parameters from high bandwidth memory (HBM) to cache. On the other…

Machine Learning · Computer Science 2024-02-15 Yashas Samaga B L , Varun Yerram , Chong You , Srinadh Bhojanapalli , Sanjiv Kumar , Prateek Jain , Praneeth Netrapalli

Executing machine learning inference tasks on resource-constrained edge devices requires careful hardware-software co-design optimizations. Recent examples have shown how transformer-based deep neural network models such as ALBERT can be…

Machine Learning · Computer Science 2023-04-14 Zirui Fu , Aleksandre Avaliani , Marco Donato

We propose to focus on the problem of discovering neural network architectures efficient in terms of both prediction quality and cost. For instance, our approach is able to solve the following tasks: learn a neural network able to predict…

Machine Learning · Computer Science 2018-05-24 Tom Veniat , Ludovic Denoyer

Network binarization is a promising hardware-aware direction for creating efficient deep models. Despite its memory and computational advantages, reducing the accuracy gap between binary models and their real-valued counterparts remains an…

Computer Vision and Pattern Recognition · Computer Science 2021-04-01 Adrian Bulat , Brais Martinez , Georgios Tzimiropoulos

Designing and implementing efficient, provably correct parallel machine learning (ML) algorithms is challenging. Existing high-level parallel abstractions like MapReduce are insufficiently expressive while low-level tools like MPI and…

Machine Learning · Computer Science 2010-06-28 Yucheng Low , Joseph Gonzalez , Aapo Kyrola , Danny Bickson , Carlos Guestrin , Joseph M. Hellerstein

Designing and implementing efficient, provably correct parallel machine learning (ML) algorithms is challenging. Existing high-level parallel abstractions like MapReduce are insufficiently expressive while low-level tools like MPI and…

Machine Learning · Computer Science 2014-08-12 Yucheng Low , Joseph E. Gonzalez , Aapo Kyrola , Danny Bickson , Carlos E. Guestrin , Joseph Hellerstein

Machine Learning algorithms based on Brain-inspired Hyperdimensional(HD) computing imitate cognition by exploiting statistical properties of high-dimensional vector spaces. It is a promising solution for achieving high energy efficiency in…

Machine Learning · Computer Science 2022-10-12 Samuel Bosch , Alexander Sanchez de la Cerda , Mohsen Imani , Tajana Simunic Rosing , Giovanni De Micheli

Most uses of machine learning today involve training a model from scratch for a particular task, or sometimes starting with a model pretrained on a related task and then fine-tuning on a downstream task. Both approaches offer limited…

Machine Learning · Computer Science 2022-05-26 Andrea Gesmundo , Jeff Dean

Binary Neural Networks (BNNs), which constrain both weights and activations to binary values, offer substantial reductions in computational complexity, memory footprint, and energy consumption. These advantages make them particularly well…

Machine Learning · Computer Science 2026-02-18 Luca Colombo , Fabrizio Pittorino , Daniele Zambon , Carlo Baldassi , Manuel Roveri , Cesare Alippi

Adept network management is key for supporting extremely heterogeneous applications with stringent quality of service (QoS) requirements; this is more so when envisioning the complex and ultra-dense 6G mobile heterogeneous network (HetNet).…

Networking and Internet Architecture · Computer Science 2022-07-04 Mohammad Arif Hossain , Abdullah Ridwan Hossain , Nirwan Ansari

Bayesian neural networks (BNNs) demonstrate promising success in improving the robustness and uncertainty quantification of modern deep learning. However, they generally struggle with underfitting at scale and parameter efficiency. On the…

In modern computer architectures, the performance of many memory-bound workloads (e.g., machine learning, graph processing, databases) is limited by the data movement bottleneck that emerges when transferring large amounts of data between…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-08-12 Pedro Carrinho , Hamid Moghadaspour , Oscar Ferraz , João Dinis Ferreira , Yann Falevoz , Vitor Silva , Gabriel Falcao

Bayesian neural networks (BNNs) have been long considered an ideal, yet unscalable solution for improving the robustness and the predictive uncertainty of deep neural networks. While they could capture more accurately the posterior…

Computer Vision and Pattern Recognition · Computer Science 2021-03-26 Gianni Franchi , Andrei Bursuc , Emanuel Aldea , Severine Dubuisson , Isabelle Bloch

The next generation of particle physics experiments will face a new era of challenges in data acquisition, due to unprecedented data rates and volumes along with extreme environments and operational constraints. Harnessing this data for…

Instrumentation and Detectors · Physics 2026-03-12 Julia Gonski , Jenni Ott , Shiva Abbaszadeh , Sagar Addepalli , Matteo Cremonesi , Jennet Dickinson , Giuseppe Di Guglielmo , Erdem Yigit Ertorer , Lindsey Gray , Ryan Herbst , Christian Herwig , Tae Min Hong , Benedikt Maier , Maryam Bayat Makou , David Miller , Mark S. Neubauer , Cristián Peña , Dylan Rankin , Seon-Hee , Seo , Giordon Stark , Alexander Tapper , Audrey Corbeil Therrien , Ioannis Xiotidis , Keisuke Yoshihara , G Abarajithan , Sagar Addepalli , Nural Akchurin , Carlos Argüelles , Saptaparna Bhattacharya , Lorenzo Borella , Christian Boutan , Tom Braine , James Brau , Martin Breidenbach , Antonio Chahine , Talal Ahmed Chowdhury , Yuan-Tang Chou , Seokju Chung , Alberto Coppi , Mariarosaria D'Alfonso , Abhilasha Dave , Chance Desmet , Angela Di Fulvio , Karri DiPetrillo , Javier Duarte , Auralee Edelen , Jan Eysermans , Yongbin Feng , Emmett Forrestel , Dolores Garcia , Loredana Gastaldo , Julián García Pardiñas , Lino Gerlach , Loukas Gouskos , Katya Govorkova , Carl Grace , Christopher Grant , Philip Harris , Ciaran Hasnip , Timon Heim , Abraham Holtermann , Tae Min Hong , Gian Michele Innocenti , Koji Ishidoshiro , Miaochen Jin , Jyothisraj Johnson , Stephen Jones , Andreas Jung , Georgia Karagiorgi , Ryan Kastner , Nicholas Kamp , Doojin Kim , Kyoungchul Kong , Katie Kudela , Jelena Lalic , Bo-Cheng Lai , Yun-Tsung Lai , Tommy Lam , Jeffrey Lazar , Aobo Li , Zepeng Li , Haoyun Liu , Vladimir Lončar , Luca Macchiarulo , Christopher Madrid , Benedikt Maier , Zhenghua Ma , Prashansa Mukim , Mark S. Neubauer , Victoria Nguyen , Sungbin Oh , Isobel Ojalvo , Hideyoshi Ozaki , Simone Pagan Griso , Myeonghun Park , Christoph Paus , Santosh Parajuli , Benjamin Parpillon , Sara Pozzi , Ema Puljak , Benjamin Ramhorst , Amy Roberts , Larry Ruckman , Kate Scholberg , Sebastian Schmitt , Noah Singer , Eluned Anne Smith , Alexandre Sousa , Michael Spannowsky , Sioni Summers , Yanwen Sun , Daniel Tapia Takaki , Antonino Tumeo , Caterina Vernieri , Belina von Krosigk , Yash Vora , Linyan Wan , Michael H. L. S. Wang , Amanda Weinstein , Andy White , Simon Williams , Felix Yu

Machine Learning (ML) is making a strong resurgence in tune with the massive generation of unstructured data which in turn requires massive computational resources. Due to the inherently compute- and power-intensive structure of Neural…

Machine Learning · Computer Science 2018-06-27 Behzad Salami , Osman Unsal , Adrian Cristal

With the recent growth in demand for large-scale deep neural networks, compute in-memory (CiM) has come up as a prominent solution to alleviate bandwidth and on-chip interconnect bottlenecks that constrain Von-Neuman architectures. However,…

Hardware Architecture · Computer Science 2024-03-19 Souvik Kundu , Anthony Sarah , Vinay Joshi , Om J Omer , Sreenivas Subramoney

Mixture of Experts (MoEs) have become a central component of many state-of-the-art open-source and proprietary large language models. Despite their widespread adoption, it remains unclear how close existing MoE architectures are to optimal…

Research has shown that deep neural networks contain significant redundancy, and that high classification accuracies can be achieved even when weights and activations are quantised down to binary values. Network binarisation on FPGAs…

Machine Learning · Computer Science 2019-04-02 Erwei Wang , James J. Davis , Peter Y. K. Cheung , George A. Constantinides

With promising yet saturated results in high-resource settings, low-resource datasets have gradually become popular benchmarks for evaluating the learning ability of advanced neural networks (e.g., BigBench, superGLUE). Some models even…

Computation and Language · Computer Science 2023-03-10 Yudong Wang , Chang Ma , Qingxiu Dong , Lingpeng Kong , Jingjing Xu

Multi-bit quantization networks enable flexible deployment of deep neural networks by supporting multiple precision levels within a single model. However, existing approaches suffer from significant training overhead as full-dataset updates…

Computer Vision and Pattern Recognition · Computer Science 2025-10-24 Jinhee Kim , Jae Jun An , Kang Eun Jeon , Jong Hwan Ko