Related papers: Ultrafast Non-Volatile Weyl LuminoMem for Mid-Infr…
We present VitaLLM, a mixed precision accelerator that enables ternary weight large language models to run efficiently on edge devices. The design combines two compute cores, a multiplier free TINT core for ternary-INT projections and a…
Zero-shot object detection enables recognising novel objects without task-specific training, but current approaches rely on large vision language models (VLMs) like CLIP that require hundreds of megabytes of memory - far exceeding the…
High-performance quantum memory for quantized states of light is a prerequisite building block of quantum information technology. Despite great progresses of optical quantum memories based on interactions of light and atoms, physical…
In modern computer architectures, the performance of many memory-bound workloads (e.g., machine learning, graph processing, databases) is limited by the data movement bottleneck that emerges when transferring large amounts of data between…
The first contribution of this paper is the development of extremely dense, energy-efficient mixed-signal vector-by-matrix-multiplication (VMM) circuits based on the existing 3D-NAND flash memory blocks, without any need for their…
The demands of modern electronic components require advanced computing platforms for efficient information processing to realize in-memory operations with a high density of data storage capabilities towards developing alternatives to von…
High-performance computing underpins modern artificial intelligence (AI), enabling foundation models, real-time inference and perception in autonomous systems, and data-intensive scientific simulations. Recent advances in quantization…
We demonstrate a memory for light based on optomechanically induced transparency. We achieve a long storage time by leveraging the ultra-low dissipation of a soft-clamped mechanical membrane resonator, which oscillates at MHz frequencies.…
The stability of the integrated photonic circuits is of critical importance for many applications that require high frequency precision or robust operation over time, such as optomechanical sensing, frequency conversion, optical…
Non-Intrusive Load Monitoring (NILM) enables the disaggregation of the global power consumption of multiple loads, taken from a single smart electrical meter, into appliance-level details. State-of-the-Art approaches are based on Machine…
In this paper, we introduce LightVLM, a simple but effective method that can be seamlessly deployed upon existing Vision-Language Models (VLMs) to greatly accelerate the inference process in a training-free manner. We divide the inference…
In cloud and edge computing models, it is important that compute devices at the edge be as power efficient as possible. Long short-term memory (LSTM) neural networks have been widely used for natural language processing, time series…
Transformers have emerged as the dominant neural-network architecture, achieving state-of-the-art performance in language processing and computer vision. At the core of these models lies the attention mechanism, which requires a nonlinear,…
The rapid rise of artificial intelligence, and in-memory computing has reinvigorated research on scalable, energy-efficient, and reconfigurable photonic hardware. Non-volatile phase-change materials (PCMs) are attractive, as they offer…
Biologically-inspired computing models have made significant progress in recent years, but the conventional von Neumann architecture is inefficient for the large-scale matrix operations and massive parallelism required by these models. This…
Optical approaches have made great strides towards the goal of high-speed, energy-efficient computing necessary for modern deep learning and AI applications. Read-in and read-out of data, however, limit the overall performance of existing…
Compute-in-memory (CiM) is a promising solution for addressing the challenges of artificial intelligence (AI) and the Internet of Things (IoT) hardware such as 'memory wall' issue. Specifically, CiM employing nonvolatile memory (NVM)…
We report the experimental observation of slow-light and coherent storage in a setting where light is tightly confined in the transverse directions. By interfacing a tapered optical nanofiber with a cold atomic ensemble, electromagnetically…
Future optical quantum technologies, such as quantum networks, distributed quantum computing and sensing, demand efficient, broadband quantum memories. However, achieving high efficiency without introducing noise, reducing bandwidth, or…
The memory demands of large-scale deep neural networks (DNNs) require synaptic weight values to be stored and updated in off-chip memory like dynamic random-access memory, which reduces energy efficiency and increases training time.…