English
Related papers

Related papers: Benchmarking CRBLASTER on the 350-MHz 49-core Maes…

200 papers

Deep learning (DL) compilers rely on cost models and auto-tuning to optimize tensor programs for target hardware. However, existing approaches depend on large offline datasets, incurring high collection costs and offering suboptimal…

Machine Learning · Computer Science 2026-04-15 Chaoyao Shen , Linfeng Jiang , Yixian Shen , Tao Xu , Guoqing Li , Anuj Pathania , Andy D. Pimentel , Meng Zhang

ReRAM-based in-memory computing (IMC) architectures are promising candidates for energy-efficient matrix-vector multiplication. While scaling the size of ReRAM arrays allows for the amortization of power-hungry peripheral circuits like DACs…

Systems and Control · Electrical Eng. & Systems 2026-04-23 Ching-Yi Lin , Sahil Shah

This paper presents an automated approach for designing processors that support a subset of the RISC-V instruction set architecture (ISA) for a new class of applications at Extreme Edge. The electronics used in extreme edge applications…

Hardware Architecture · Computer Science 2025-10-29 Alireza Raisiardali , Konstantinos Iordanou , Jedrzej Kufel , Kowshik Gudimetla , Kris Myny , Emre Ozer

The aim of this study is the characterization of the computing resources used by researchers at the "Instituto de Astrof\'isica de Canarias" (IAC). Since there is a huge demand of computing time and we use tools such as HTCondor to…

Performance · Computer Science 2017-02-17 Nicola Caon , Antonio Dorta , Juan Carlos Trelles Arjona

Development of novel radiation-absorbent materials and devices for millimetre and submillimetre astronomy instruments is a research area of high interest, and with substantial engineering challenges. Alongside low-profile structure and…

Instrumentation and Methods for Astrophysics · Physics 2023-02-10 Giampaolo Pisano , Christopher Dunscombe , Peter Hargrave , Alexey Shitvov , Carole Tucker

Computation intensive kernels, such as convolutions, matrix multiplication and Fourier transform, are fundamental to edge-computing AI, signal processing and cryptographic applications. Interleaved-Multi-Threading (IMT) processor cores are…

Hardware Architecture · Computer Science 2021-02-09 Abdallah Cheikh , Stefano Sordillo , Antonio Mastrandrea , Francesco Menichelli , Giuseppe Scotti , Mauro Olivieri

This paper aims to implement the six channel redundancy to achieve fault tolerance in testing of satellites with acoustic spectrum. We mainly focus here on achieving fault tolerance. An immediate application is the microphone data…

Other Computer Science · Computer Science 2010-05-07 H. S. Aravinda , H. D. Maheshappa , Ranjan Moodithaya

In this paper, we introduce the design and verification frameworks for developing a fully-functional emerging ternary processor. Based on the existing compiling environments for binary processors, for the given ternary instructions, the…

Hardware Architecture · Computer Science 2021-11-16 Dongyun Kam , Jung Gyu Min , Jongho Yoon , Sunmean Kim , Seokhyeong Kang , Youngjoo Lee

Large-scale HPC simulations of plasma dynamics in fusion devices require efficient parallel I/O to avoid slowing down the simulation and to enable the post-processing of critical information. Such complex simulations lacking parallel I/O…

We present a first of its kind framework which overcomes a major challenge in the design of digital systems that are resilient to reliability failures: achieve desired resilience targets at minimal costs (energy, power, execution time,…

The growing demand for edge computing and AI drives research into analog in-memory computing using memristors, which overcome data movement bottlenecks by computing directly within memory. However, device failures and variations critically…

Emerging Technologies · Computer Science 2025-07-16 Zhicheng Xu , Jiawei Liu , Sitao Huang , Zefan Li , Shengbo Wang , Bo Wen , Ruibin Mao , Mingrui Jiang , Giacomo Pedretti , Jim Ignowski , Kaibin Huang , Can Li

In this thesis, we propose a light-weight sparsity-based algorithm, basic thresholding classifier (BTC), for classification applications (such as face identification, hyper-spectral image classification, etc.) which is capable of…

Computer Vision and Pattern Recognition · Computer Science 2017-12-11 Mehmet Altan Toksöz

Accurate channel modeling has become critical for evaluating multiple-input multiple-output (MIMO) performance, especially as 5G standardization matures and efforts toward 6G begin. Recent studies within the 3rd Generation Partnership…

Information Theory · Computer Science 2026-03-04 Lynda Berrah , Raphael Visoz , Didier Le Ruyet , Anvar Tukmanov , Axel Müller , Alexander Hamilton , Matthew Baker

FPGA overlays are commonly implemented as coarse-grained reconfigurable architectures with a goal to improve designers' productivity through balancing flexibility and ease of configuration of the underlying fabric. To truly facilitate full…

Hardware Architecture · Computer Science 2016-06-22 Ho-Cheung Ng , Cheng Liu , Hayden Kwok-Hay So

Compound LLM training workloads-such as knowledge distillation and multimodal LLM (MLLM) training-are gaining prominence. These typically comprise heterogeneous components differing in parameter scale, execution mode (forward-only or full…

We have integrated a system of 16 RISC CPUs to help reconstruct and analyze a 1.3 Terabyte data set of 400 million high energy physics interactions. These new CPUs provided an affordable means of processing a very large data set. The data…

High Energy Physics - Experiment · Physics 2009-10-31 C. Stoughton , D. J. Summers

As AI chips incorporate numerous parallelized cores to scale deep learning (DL) computing, inter-core communication is enabled recently by employing high-bandwidth and low-latency interconnect links on the chip (e.g., Graphcore IPU). It…

Distributed, Parallel, and Cluster Computing · Computer Science 2024-09-25 Yiqi Liu , Yuqi Xue , Yu Cheng , Lingxiao Ma , Ziming Miao , Jilong Xue , Jian Huang

Designing and validating efficient cache-coherent memory subsystems is a critical yet complex task in the development of modern multi-core system-on-chip architectures. Rhea is a unified framework that streamlines the design and…

Hardware Architecture · Computer Science 2026-03-10 Davide Zoni , Andrea Galimberti , Adriano Guarisco

In the upcoming LHC Run 3, starting in 2021, the upgraded Time Projection Chamber (TPC) of the ALICE experiment will record minimum bias Pb-Pb collisions in a continuous readout mode at an interaction rate up to 50 kHz. This corresponds to…

Instrumentation and Detectors · Physics 2021-02-03 Marten Ole Schmidt

The scientific community increasingly relies on machine learning (ML) for near-sensor processing, leveraging its strengths in tasks such as pattern recognition, anomaly detection, and real-time decision-making. These deployments demand…

Hardware Architecture · Computer Science 2026-03-30 G Abarajithan , Zhenghua Ma , Ravidu Munasinghe , Francesco Restuccia , Ryan Kastner