Related papers: SA-Kura: An Energy-Efficient Systolic Array Accele…
Score-debiased kernel density estimation (SD-KDE) achieves improved asymptotic convergence rates over classical KDE, but its use of an empirical score has made it significantly slower in practice. We show that by re-ordering the SD-KDE…
Leveraging the diffusion transformer (DiT) architecture, models like Sora, CogVideoX and Wan have achieved remarkable progress in text-to-video, image-to-video, and video editing tasks. Despite these advances, diffusion-based video…
We present a new code aimed at the simulation of diffusive shock acceleration (DSA), and discuss various test cases which demonstrate its ability to study DSA in its full time-dependent and non-linear developments. We present the numerical…
Latent Consistency Models (LCMs) have achieved impressive performance in accelerating text-to-image generative tasks, producing high-quality images with minimal inference steps. LCMs are distilled from pre-trained latent diffusion models…
Access to 3D point cloud representations has been widely facilitated by LiDAR sensors embedded in various mobile devices. This has led to an emerging need for fast and accurate point cloud processing techniques. In this paper, we revisit…
Cornell's electron/positron storage ring (CESR) was modified over a series of accelerator shutdowns beginning in May 2008, which substantially improves its capability for research and development for particle accelerators. CESR's energy…
Transformers achieve state-of-the-art performance in natural language processing, vision, and scientific computing, but demand high computation and memory. To address these challenges, we present ASTRA, the first silicon-photonic…
The next generation HPC and data centers are likely to be reconfigurable and data-centric due to the trend of hardware specialization and the emergence of data-driven applications. In this paper, we propose ARENA -- an asynchronous…
This work includes all the technical details of the Sequential Principal Curves Analysis (SPCA) in a single document. SPCA is an unsupervised nonlinear and invertible feature extraction technique. The identified curvilinear features can be…
Memristive crossbars have become a popular means for realizing unsupervised and supervised learning techniques. In previous neuromorphic architectures with leaky integrate-and-fire neurons, the crossbar itself has been separated from the…
Charged particle accelerators play a pivotal role in scientific research, industry, and medical applications. Among them, radiofrequency (RF) accelerators offer a promising approach for achieving high-energy particle acceleration in compact…
We demonstrate a compact technique to compress electron pulses to attosecond length, while keeping the energy spread reasonably small. The technique is based on Dielectric Laser Acceleration (DLA) in nanophotonic silicon structures. Unlike…
Single-issue processor cores are very energy efficient but suffer from the von Neumann bottleneck, in that they must explicitly fetch and issue the loads/storse necessary to feed their ALU/FPU. Each instruction spent on moving data is a…
The number of processing elements (PEs) in a fixed-sized systolic accelerator is well matched for large and compute-bound DNNs; whereas, memory-bound DNNs suffer from PE underutilization and fail to achieve peak performance and energy…
Traditional super-resolution (SR) methods assume an ``ideal'' downscaling SR-kernel (e.g., bicubic downscaling) between the high-resolution (HR) image and the low-resolution (LR) image. Such methods fail once the LR images are generated…
We propose a flux-pumped superconducting parametric amplifier based on symmetrically threaded superconducting quantum interference devices (SQUIDs) that achieves a Kerr-free operating point under suitable drive conditions. Eliminating the…
The influence of array geometry on synchronization properties of a 2-D oscillator array is investigated based on a comparison between a rectangular and a hexagonal grid. The Kuramoto model is solved for a nearest-neighbor case with periodic…
Current 3D single object tracking methods primarily rely on the Siamese matching-based paradigm, which struggles with textureless and incomplete LiDAR point clouds. Conversely, the motion-centric paradigm avoids appearance matching, thus…
It has long been known in both neuroscience and AI that ``binding'' between neurons leads to a form of competitive learning where representations are compressed in order to represent more abstract concepts in deeper layers of the network.…
The steeply growing performance demands for highly power- and energy-constrained processing systems such as end-nodes of the internet-of-things (IoT) have led to parallel near-threshold computing (NTC), joining the energy-efficiency…