Related papers: Fine Grain 3D Integration for Microarchitecture De…
We present a fabrication process for fully superconducting interconnects compatible with superconducting qubit technology. These interconnects allow for the 3D integration of quantum circuits without introducing lossy amorphous dielectrics.…
Polyurethane (PU) possesses excellent thermal properties, making it an ideal material for thermal insulation. Incorporating Phase Change Materials (PCMs) capsules into Polyurethane (PU) has proven to be an effective strategy for enhancing…
Photonic Integrated Circuits (PICs) provide superior speed, bandwidth, and energy efficiency, making them ideal for communication, sensing, and quantum computing applications. Despite their potential, PIC design workflows and integration…
When partitioning gate-level netlists using graphs, it is beneficial to cluster gates to reduce the order of the graph and preserve some characteristics of the circuit that the partitioning might degrade. Gate clustering is even more…
Design of printed circuit board (PCB) stack-up requires the consideration of characteristic impedance, insertion loss and crosstalk. As there are many parameters in a PCB stack-up design, the optimization of these parameters needs to be…
Silicon carbide (SiC) is an important semiconductor material for fabricating power electronic devices that exhibit higher switch frequency, lower energy loss and substantial reduction both in size and weight in comparison with its Si-based…
Large language model (LLM) decoding is a major inference bottleneck because its low arithmetic intensity makes performance highly sensitive to memory bandwidth. 3D-stacked near-memory processing (NMP) provides substantially higher local…
Recent advances in three-dimensional laser writing have enabled direct nanostructuring deep within silicon, unlocking a volumetric design space previously inaccessible to surface-bound nanophotonic devices. Here, we introduce subwavelength…
Modular trapped-ion quantum computing hardware, known as QCCDs require shuttling operations in order to maintain effective all-to-all connectivity. Each module or trap can perform only one operation at a time, resulting in low intra-trap…
Analytical hardware performance models yield swift estimation of desired hardware performance metrics. However, developing these analytical models for modern processors with sophisticated microarchitectures is an extremely laborious task…
In this paper we discuss design concepts for increasing the spatial resolution, improving the sensitivity, and reducing the invasiveness in scanning Superconducting Quantum Interference Device (SQUID) microscope sensors with integrated flux…
Today's computing systems require moving data back-and-forth between computing resources (e.g., CPUs, GPUs, accelerators) and off-chip main memory so that computation can take place on the data. Unfortunately, this data movement is a major…
Bloom filters are a fundamental data structure for approximate membership queries, with applications ranging from data analytics to databases and genomics. Several variants have been proposed to accommodate parallel architectures. GPUs,…
Cosmic dust particles effectively attenuate starlight. Their absorption of starlight produces emission spectra from the near- to far-infrared, which depends on the sizes and properties of the dust grains, and spectrum of the heating…
This paper proposes a new parallel approach to solve connected components on a 2D binary image implemented with CUDA. We employ the following strategies to accelerate neighborhood exploration after dividing an input image into independent…
We investigate the role of microstructural bridging on the fracture toughness of composite materials. To achieve this, a new computational framework is presented that integrates phase field fracture and cohesive zone models to simulate…
The unabated growth in AI workload demands is driving the need for concerted advances in compute, memory, and interconnect performance. As traditional semiconductor scaling slows, high-speed interconnects have emerged as the new scaling…
We propose an optimization approach for determining both hardware and software parameters for the efficient implementation of a (family of) applications called dense stencil computations on programmable GPGPUs. We first introduce a simple,…
Transformer models have revolutionized AI tasks, but their large size hinders real-world deployment on resource-constrained and latency-critical edge devices. While binarized Transformers offer a promising solution by significantly reducing…
This paper presents a comprehensive study of interactive rendering techniques for large 3D line sets with transparency. The rendering of transparent lines is widely used for visualizing trajectories of tracer particles in flow fields.…