English
Related papers

Related papers: Performance Analysis of Edge and In-Sensor AI Proc…

200 papers

Wearable biosignal processing applications are driving significant progress toward miniaturized, energy-efficient Internet-of-Things solutions for both clinical and consumer applications. However, scaling toward high-density multi-channel…

Systems and Control · Electrical Eng. & Systems 2023-07-06 Sebastian Frey , Marco Guermandi , Simone Benatti , Victor Kartsch , Andrea Cossettini , Luca Benini

Modern exascale GPU- and APU-based systems provide multiple power and energy sensors, but differences in scope, update rate, timing, and filtering complicate the attribution of short-lived accelerator activity. This paper presents a…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-04-13 Adam McDaniel , Michael Jantz , Ashesh Sharma , Steve Abbott , Steven Martin , Shreyas Khandekar , Brandon Neth , Bruno Villasenor Alvarez , Aditya Kashi , Wael Elwasif , Oscar Hernandez

The rapid adaptation of data driven AI models, such as deep learning inference, training, Vision Transformers (ViTs), and other HPC applications, drives a strong need for runtime precision configurable different non linear activation…

Hardware Architecture · Computer Science 2026-02-12 Mukul Lokhande , Gopal Raut , Santosh Kumar Vishvakarma

This review explores the intersection of bio-plausible artificial intelligence in the form of Spiking Neural Networks (SNNs) with the analog In-Memory Computing (IMC) domain, highlighting their collective potential for low-power edge…

Neural and Evolutionary Computing · Computer Science 2024-09-20 Abhishek Moitra , Abhiroop Bhattacharjee , Yuhang Li , Youngeun Kim , Priyadarshini Panda

The rapid development in scientific research provides a need for more compute power, which is partly being solved by GPUs. This paper presents a microarchitectural analysis of the modern NVIDIA Blackwell architecture by studying GPU…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-07-23 Aaron Jarmusch , Nathan Graddon , Sunita Chandrasekaran

Computing-in-memory (CIM) is an emerging computing paradigm, offering noteworthy potential for accelerating neural networks with high parallelism, low latency, and energy efficiency compared to conventional von Neumann architectures.…

Neural and Evolutionary Computing · Computer Science 2024-09-30 Kam Chi Loong , Shihao Han , Sishuo Liu , Ning Lin , Zhongrui Wang

We present a scalable in-pixel processing architecture that can reduce the data throughput by 10X and consume less than 30 mW per megapixel at the imager frontend. Unlike the state-of-the-art (SOA) analog process-in-pixel (PIP) that…

Hardware Architecture · Computer Science 2022-10-17 David Zhang , Gooitzen van der Wal , Saurabh Farkya , Thomas Senko , Aswin Raghavan , Michael Isnardi , Michael Piacentino

As cost and performance benefits associated with Moore's Law scaling slow, researchers are studying alternative architectures (e.g., based on analog and/or spiking circuits) and/or computational models (e.g., convolutional and recurrent…

Emerging Technologies · Computer Science 2019-06-14 Qiuwen Lou , Indranil Palit , Tang Li , Andras Horvath , Michael Niemier , X. Sharon Hu

Edge-AI applications still face considerable challenges in enhancing computational efficiency in resource-constrained environments. This work presents RAMAN, a resource-efficient and approximate posit(8,2)-based Multiply-Accumulate (MAC)…

Hardware Architecture · Computer Science 2025-10-28 Mohd Faisal Khan , Mukul Lokhande , Santosh Kumar Vishvakarma

The performance of mobile AI accelerators has been evolving rapidly in the past two years, nearly doubling with each new generation of SoCs. The current 4th generation of mobile NPUs is already approaching the results of CUDA-compatible…

Performance · Computer Science 2019-10-16 Andrey Ignatov , Radu Timofte , Andrei Kulik , Seungsoo Yang , Ke Wang , Felix Baum , Max Wu , Lirong Xu , Luc Van Gool

Generative AI (GenAI) services powered by large language models (LLMs) increasingly deliver real-time interactions, yet existing 5G multi-access edge computing (MEC) architectures often treat communication and computing as separate domains,…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-04-24 Chien-Sheng Yang , Yu-Jen Ku , Yuan-Yao Lou , Nathan Tenny , Alex C. -C. Hsu

Edge-device co-inference refers to deploying well-trained artificial intelligent (AI) models at the network edge under the cooperation of devices and edge servers for providing ambient intelligent services. For enhancing the utilization of…

Information Theory · Computer Science 2023-08-15 Zeming Zhuang , Dingzhu Wen , Yuanming Shi , Guangxu Zhu , Sheng Wu , Dusit Niyato

Edge computing is a promising solution for handling high-dimensional, multispectral analog data from sensors and IoT devices for applications such as autonomous drones. However, edge devices' limited storage and computing resources make it…

Machine Learning · Computer Science 2023-09-21 Nastaran Darabi , Amit R. Trivedi

Huge energy consumption poses a significant challenge for edge clouds. In response to this, we introduce a new type of edge server, namely SoC Cluster, that orchestrates multiple low-power mobile system-on-chips (SoCs) through an on-chip…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-07-18 Li Zhang , Zhe Fu , Boqing Shi , Xiang Li , Rujin Lai , Chenyang Yang , Ao Zhou , Xiao Ma , Shangguang Wang , Mengwei Xu

The deployment of transformer-based models on resource-constrained edge devices represents a critical challenge in enabling real-time artificial intelligence applications. This comprehensive survey examines lightweight transformer…

Machine Learning · Computer Science 2026-01-08 Hema Hariharan Samson

Artificial intelligence and machine learning models deployed on edge devices, e.g., for quality control in Additive Manufacturing (AM), are frequently small in size. Such models usually have to deliver highly accurate results within a short…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-11-26 Marcel Aach , Cyril Blanc , Andreas Lintermann , Kurt De Grave

AI models are increasing in size and recent advancement in the community has shown that unlike HPC applications where double precision datatype are required, lower-precision datatypes such as fp8 or int4 are sufficient to bring the same…

Performance · Computer Science 2023-10-11 Saeed Maleki

This paper presents a novel approach to event-based power modelling for embedded platforms that do not have a Performance Monitoring Unit (PMU). The method involves complementing the target hardware platform, where the physical power data…

Distributed, Parallel, and Cluster Computing · Computer Science 2022-02-07 Kris Nikov , Marcos Martinez , Simon Wegener , Jose Nunez-Yanez , Zbigniew Chamski , Kyriakos Georgiou , Kerstin Eder

The rapid scaling of large language models (LLMs) has unveiled critical limitations in current hardware architectures, including constraints in memory capacity, computational efficiency, and interconnection bandwidth. DeepSeek-V3, trained…

Distributed, Parallel, and Cluster Computing · Computer Science 2025-12-24 Chenggang Zhao , Chengqi Deng , Chong Ruan , Damai Dai , Huazuo Gao , Jiashi Li , Liyue Zhang , Panpan Huang , Shangyan Zhou , Shirong Ma , Wenfeng Liang , Ying He , Yuqing Wang , Yuxuan Liu , Y. X. Wei

Spiking Neural Networks (SNNs) have gained significant attention in edge computing due to their low power consumption and computational efficiency. However, existing implementations either use conventional System on Chip (SoC) architectures…

Hardware Architecture · Computer Science 2026-03-13 Kanishka Gunawardana , Sanka Peeris , Kavishka Rambukwella , Thamish Wanduragala , Saadia Jameel , Roshan Ragel , Isuru Nawinne