Related papers: Nonvolatile Charge-Domain Attention with HZO Ferro…
In this work, we report a novel design, one-transistor-one-inverter (1T1I), to satisfy high speed and low power on-chip training requirements. By leveraging doped HfO2 with ferroelectricity, a non-volatile inverter is successfully…
Gated DeltaNet (GDN) is a linear attention mechanism that replaces the growing KV cache with a fixed-size recurrent state. Hybrid LLMs like Qwen3-Next use 75% GDN layers and achieve competitive accuracy to attention-only models. However, at…
We present a flexible, automated, and basis-set insensitive domain-based charge-transfer (CT) decomposition framework that can be combined with any CI-type excited-state wavefunction. Our approach is not based on excited-state densities and…
Multimodal large language models suffer from substantial inference overhead since multimodal KV Cache grows proportionally with the visual input length. Existing multimodal KV Cache compression methods mostly rely on attention score to…
In this study, we show that the discharge voltage pattern of a fractional-order supercapacitor from the same initial steady-state voltage into a constant resistor is dependent on the past charging voltage profile. The charging voltage was…
We investigate in detail antiferromagnetic (AF) and superconducting (SC) phases as well as their coexistence in the two-dimensional Kondo lattice model on a square lattice, which is a paradigmatic model for heavy fermion materials. The…
We report the site-specific probing of charge-transfer dynamics in a prototype system for organic photovoltaics (OPV) by picosecond time-resolved X-ray photoelectron spectroscopy. A layered system consisting of approximately two monolayers…
PCM is a popular backing memory for DRAM main memory in tiered memory systems. PCM has asymmetric access energy; writes dominate reads. MLC asymmetry can vary by an order of magnitude. Many schemes have been developed to take advantage of…
This work presents a novel approach to configure 2T-nC ferroelectric RAM (FeRAM) for performing single cell logic-in-memory operations, highlighting its advantages in energy-efficient computation over conventional DRAM-based approaches.…
Quantum reservoir computing (QRC) offers a promising framework for online quantum-enhanced machine learning tailored to temporal tasks, yet practical implementations with native memory capabilities remain limited. Here, we demonstrate an…
This paper presents a memory assessment of the next-generation Versatile Video Coding (VVC). The memory analyses are performed adopting as a baseline the state-of-the-art High-Efficiency Video Coding (HEVC). The goal is to offer insights…
Driver vigilance estimation is an important task for transportation safety. Wearable and portable brain-computer interface devices provide a powerful means for real-time monitoring of the vigilance level of drivers to help with avoiding…
Transition metal oxides (TMOs) and post-TMOs (PTMOs), when doped with Carbon, show non-volatile current-voltage (I-V) characteristics, which are both universal and repeatable. We have shown spectroscopic evidence of the introduction of…
Robust multi-level spin memory with the ability to write information electrically is a long-sought capability in spintronics, with great promise for applications. Here we achieve nonvolatile and highly energy-efficient magnetization…
Long-horizon LLM inference turns the key--value (KV) cache into the dominant GPU memory consumer and makes per-token attention increasingly expensive. Many common eviction policies use static recency windows or historical attention, leaving…
We extend the scope of full configuration interaction quantum Monte Carlo (FCIQMC) to be applied to coupled fermion-boson hamiltonians, alleviating the a priori truncation in boson occupation which is necessary for many other wave function…
Multi-head latent attention (MLA) is designed to optimize KV cache memory through low-rank key-value joint compression. Rather than caching keys and values separately, MLA stores their compressed latent representations, reducing memory…
Compute-in-memory (CIM) presents an attractive approach for energy-efficient computing in data-intensive applications. However, the development of suitable memory designs to achieve high-performance CIM remains a challenging task. Here, we…
Transformer inference requires high compute accuracy; achieving this using analog CIMs has been difficult due to inherent computational errors. To overcome this challenge, we propose a Capacitor-Reconfiguring CIM (CR-CIM) to realize high…
The rapid growth of digital technology has driven the need for efficient storage solutions, positioning memristors as promising candidates for next-generation non-volatile memory (NVM) due to their superior electrical properties. Organic…