Related papers: Taming Process Variations in CNFET for Efficient L…
This study presents a methodology for anticounterfeiting of Non-Volatile Memory (NVM) chips. In particular, we experimentally demonstrate a generalized methodology for detecting (i) Integrated Circuit (IC) origin, (ii) recycled or used NVM…
Fabrics of continuous fibers of carbon nanotubes (CNTFs) are attractive materials for multifunctional energy storage devices, either as current collector, or as active material. Despite a similar chemical composition,…
Spin-Transfer Torque Magnetic RAM (STT-MRAM) is known as the most promising replacement for SRAM technology in large Last-Level Caches (LLCs). Despite its high-density, non-volatility, near-zero leakage power, and immunity to radiation as…
Using accurate Hybrid-Functional DFT coupled with the Non-Equilibrium Green's function (NEGF) formalism, we explore and benchmark the fundamental scaling limits of CNT-FETs against Si and 2D-material MoS$_2$ and HfS$_2$ Nanosheets,…
A physics-based compact model for silicon gate-all-around (GAA) nanowire tunneling FETs (NW-tFETs) with good accuracy has been developed by considering Phonon-Assisted Tunneling (PAT) and transition from Quantum Capacitance Limit (QCL) to…
The end of Dennard scaling has pushed power consumption into a first order concern for current systems, on par with performance. As a result, near-threshold voltage computing (NTVC) has been proposed as a potential means to tackle the…
We demonstrate solution processable large area field effect transistors (FETs) from aligned arrays of carbon nanotubes (CNTs). Commercially available, surfactant free CNTs suspended in aqueous solution were aligned between source and drain…
This paper investigates the usage of FPGA devices for energy-efficient exact kNN search in high-dimension latent spaces. This work intercepts a relevant trend that tries to support the increasing popularity of learned representations based…
The rapid growth of LLMs demands high-throughput, memory-capacity-intensive inference on resource-constrained edge devices, where single-batch decoding remains fundamentally memory-bound. Existing out-of-core GPU-based and SSD-like…
Multicore processors constitute the main architecture choice for modern computing systems in different market segments. Despite their benefits, the contention that naturally appears when multiple applications compete for the use of shared…
Layered Li(Ni,Mn,Co,)O$_2$ (NMC) presents an intriguing ternary alloy design space for optimization as a cathode material in Li-ion batteries. Recently, the high cost and resource limitations of Co have added a new design constraint and…
Convolutional Neural Networks (CNNs) are widely used in deep learning applications, e.g. visual systems, robotics etc. However, existing software solutions are not efficient. Therefore, many hardware accelerators have been proposed…
This paper presents Flash, an optimized private inference (PI) hybrid protocol utilizing both homomorphic encryption (HE) and secure two-party computation (2PC), which can reduce the end-to-end PI latency for deep CNN models less than 1…
Efficient deployment of neural networks (NN) requires the co-optimization of accuracy and latency. For example, hardware-aware neural architecture search has been used to automatically find NN architectures that satisfy a latency constraint…
Among hardware accelerators for deep-learning inference, data flow implementations offer low latency and high throughput capabilities. In these architectures, each neuron is mapped to a dedicated hardware unit, making them well-suited for…
The material and electrical properties of the CNT single vias and array vias grown by microwave plasma-enhanced chemical vapor deposition were investigated. The diameters of multiwall carbon nanotubes (MWNTs) grown on the bottom electrode…
Feature caching approaches accelerate diffusion transformers (DiTs) by storing the output features of computationally expensive modules at certain timesteps, and exploiting them for subsequent steps to reduce redundant computations. Recent…
In recent years, neural image compression (NIC) algorithms have shown powerful coding performance. However, most of them are not adaptive to the image content. Although several content adaptive methods have been proposed by updating the…
Multi-headed Attention's (MHA) quadratic compute and linearly growing KV-cache make long-context transformers expensive to train and serve. Prior works such as Grouped Query Attention (GQA) and Multi-Latent Attention (MLA) shrink the cache,…
Learned image compression (LIC) methods have exhibited promising progress and superior rate-distortion performance compared with classical image compression standards. Most existing LIC methods are Convolutional Neural Networks-based…