Related papers: Integrated electro-optic attention nonlinearities …
The transformer structure employed in large language models (LLMs), as a specialized category of deep neural networks (DNNs) featuring attention mechanisms, stands out for their ability to identify and highlight the most relevant aspects of…
The rapid growth of large-scale AI models has intensified energy consumption and data-movement challenges in modern datacenters. Photonic accelerators offer a promising path by executing the linear matrix multiplications of transformer…
Despite their central role in the success of foundational models and large-scale language modeling, the theoretical foundations governing the operation of Transformers remain only partially understood. Contemporary research has largely…
Integrated photonic platforms can greatly enhance the efficiency of nonlinear frequency conversion processes by tightly confining light on a sub-micron scale. However, this advantage is often reduced by large fiber-to-chip coupling losses…
Optical computing offers potential for ultra high-speed and low latency computation by leveraging the intrinsic properties of light. Here, we explore the use of highly nonlinear optical fibers (HNLFs) as platforms for optical computing…
The Transformer model has been pivotal in advancing fields such as natural language processing, speech recognition, and computer vision. However, a critical limitation of this model is its quadratic computational and memory complexity…
Widely adopted in modern Vision Transformer designs, Softmax attention can effectively capture long-range visual information; however, it incurs excessive computational cost when dealing with high-resolution inputs. In contrast, linear…
Recent advances in transformer-based Large Language Models (LLMs) have demonstrated remarkable capabilities across various tasks. However, their quadratic computational complexity concerning sequence length remains a significant bottleneck…
Transformer-based models have emerged as one of the most widely used architectures for natural language processing, natural language generation, and image generation. The size of the state-of-the-art models has increased steadily reaching…
Electro-optic modulation performs a technological relevant functionality such as for communication, beam steering, or neuromorphic computing through providing the nonlinear activation function of a perceptron. Wile Silicon photonics enabled…
Neural operators offer a powerful data-driven framework for learning mappings between function spaces, in which the transformer-based neural operator architecture faces a fundamental scalability-accuracy trade-off: softmax attention…
Metasurfaces represent a pivotal advancement in nonlinear optics, leveraging high-Q resonant cavities to enhance harmonic generation. Multi-layer metasurfaces (MLMs) further amplify this potential by intensifying light-matter interactions…
Nonlinear imaging systems can surpass the limits of linear optics, but to date they have all relied on physical media (e.g. crystals) to work. These materials are all constrained by their physical properties, such as frequency selectivity,…
Transformer architectures based on the attention mechanism have revolutionized natural language processing (NLP), driving major breakthroughs across virtually every NLP task. However, their substantial memory and computational requirements…
Although transformers have become the neural architectures of choice for natural language processing, they require orders of magnitude more training data, GPU memory, and computations in order to compete with convolutional neural networks…
We study architectural and optimization techniques for sample-efficient language modeling under the constraints of the BabyLM 2025 shared task. Our model, BLaLM, replaces self-attention with a linear-time mLSTM token mixer and explores…
Diamond photonics has enabled efficient interfaces for quantum memories and is predicted to be a critical component of quantum networks. However, scalable network architectures require spatial, temporal, and spectral control of photons,…
Electro-optic modulators provide a key function in optical transceivers and increasingly in photonic programmable Application Specific Integrated Circuits (ASICs) for machine learning and signal processing. However, both foundry ready…
Self-attention mechanism is the key of the Transformer but often criticized for its computation demands. Previous token pruning works motivate their methods from the view of computation redundancy but still need to load the full network and…
The integrated photonics CMOS-compatible silicon nitride (SiN) platform is praised for its low propagation loss, but is limited by its lack of active functionalities such as a strong Pockels coefficient and intrinsic \c{hi}(2) nonlinearity.…