Related papers: System-Technology Co-Optimization of Bitline Routi…
Medium-scale or large-scale receive antenna array with digital beamforming can be employed at receiver to make a significant interference reduction, but leads to expensive cost and high complexity of the RF-chain circuit. To deal with this…
Super-resolution ultrasound (SRUS) imaging through localising and tracking sparse microbubbles has been shown to reveal microvascular structure and flow beyond the wave diffraction limit. Most SRUS studies use standard delay and sum (DAS)…
This paper summarizes the idea of Tiered-Latency DRAM (TL-DRAM), which was published in HPCA 2013, and examines the work's significance and future potential. The capacity and cost-per-bit of DRAM have historically scaled to satisfy the…
Deep neural networks are widely deployed in many fields. Due to the in-situ computation (known as processing in memory) capacity of the Resistive Random Access Memory (ReRAM) crossbar, ReRAM-based accelerator shows potential in accelerating…
This paper investigates hardware-based memory compression designs to increase the memory bandwidth. When lines are compressible, the hardware can store multiple lines in a single memory location, and retrieve all these lines in a single…
Fast directional solidification during Laser Additive Manufacturing (LAM) produces a complex microstructure in nickel-based superalloys, comprising columnar grains with cellular sub-grain structures and carbides. Using non-destructive…
To continue scaling beyond 2-D CMOS with 3-D integration, any new 3-D IC technology has to be comparable or better than 2-D CMOS in terms of scalability, enhanced functionality, density, power, performance, cost, and reliability.…
The human brain simultaneously optimizes synaptic weights and topology by growing, pruning, and strengthening synapses while performing all computation entirely in memory. In contrast, modern artificial-intelligence systems separate weight…
This paper investigates bandwidth-efficient DRAM caching for hybrid DRAM + 3D-XPoint memories. 3D-XPoint is becoming a viable alternative to DRAM as it enables high-capacity and non-volatile main memory systems; however, 3D-XPoint has 4-8x…
We propose a new signaling scheme for on-chip optical-electrical-optical artificial neural networks that utilizes orthogonal delay-division multiplexing and pilot-tone based self-homodyne detection. This scheme offers a more efficient…
Hybrid neural ordinary differential equations (neural ODEs) integrate mechanistic models with neural ODEs, offering strong inductive bias and flexibility, and are particularly advantageous in data-scarce healthcare settings. However,…
We employ accurate density functional theory calculations to examine the electronic structure of three Ni/SrO$_2$ nanostructures containing single-layer, bilayer and four-layer Ni nanosheets. The single Ni layer interacts strongly with the…
High drive current is a critical performance parameter in semiconductor devices for high-speed, low-power logic applications or high-efficiency, high-power, high-speed radio frequency (RF) analog applications. In this work, we demonstrate…
Bit-serial architectures can handle Neural Networks (NNs) with different weight precisions, achieving higher resource efficiency compared with bit-parallel architectures. Besides, the weights contain abundant zero bits owing to the fault…
We investigate a reconfigurable intelligent surface (RIS)-aided multi-user massive multiple-input multi-output (MIMO) system where low-resolution digital-analog converters (DACs) are configured at the base station (BS) in order to reduce…
Machine learning training places immense demands on cluster networks, motivating specialized architectures and co-design with parallelization strategies. Recent designs incorporating optical circuit switches (OCSes) are promising, offering…
Many quantum computing platforms are based on a two-dimensional physical layout. Here we explore a concept called looped pipelines which permits one to obtain many of the advantages of a 3D lattice while operating a strictly 2D device. The…
Monolithic 3D (M3D) technology enables high density integration, performance, and energy-efficiency by sequentially stacking tiers on top of each other. M3D-based network-on-chip (NoC) architectures can exploit these benefits by adopting…
Conventional 6T SRAM is used in microprocessors in the cache memory design. The basic 6T SRAM cell and a 6 bit memory array layout are designed in LEdit. The design and analysis of key SRAM components, sense amplifiers, decoders, write…
To support the boosting interconnect capacity of the AI-related data centers, novel techniques enabled high-speed and low-cost optics are continuously emerging. When the baud rate approaches 200 GBaud per lane, the bottle-neck of…