相关论文: Self-attention-based non-linear basis transformati…
Lattices are an efficient and effective method to encode ambiguity of upstream systems in natural language processing tasks, for example to compactly capture multiple speech recognition hypotheses, or to represent multiple linguistic…
The ability to form images through hair-thin optical fibres promises to open up new applications from biomedical imaging to industrial inspection. Unfortunately, deployment has been limited because small changes in mechanical deformation…
While the Self-Attention mechanism in the Transformer model has proven to be effective in many domains, we observe that it is less effective in more diverse settings (e.g. multimodality) due to the varying granularity of each token and the…
Ultra-thin multimode optical fiber imaging promises next-generation medical endoscopes reaching high image resolution for deep tissues. However, current technology suffers from severe optical distortion, as the fiber's calibration is…
When light propagates through a complex medium, such as a multimode optical fibre (MMF), the spatial information it carries is scrambled. In this work we experimentally demonstrate an all-optical strategy to unscramble this light again. We…
Transformers have become one of the dominant architectures in deep learning, particularly as a powerful alternative to convolutional neural networks (CNNs) in computer vision. However, Transformer training and inference in previous works…
When multimode optical fibers are perturbed, the data that is transmitted through them is scrambled. This presents a major difficulty for many possible applications, such as multimode fiber-based telecommunication and endoscopy. To overcome…
Recurrent neural networks excel at temporal tasks and video processing but require energy-intensive sequential memory operations. We demonstrate that multimode optical fibers naturally implement spatiotemporal recurrent computation through…
When light propagates through opaque material, the spatial information it holds becomes scrambled, but not necessarily lost. Two classes of techniques have emerged to recover this information: methods relying on optical memory effects, and…
When light propagates through a multimode optical fibre (MMF), the spatial information it carries is scrambled. Wavefront shaping can undo this scrambling, typically one spatial mode at a time - enabling deployment of MMFs as ultra-thin…
Attention mechanisms, especially self-attention, have played an increasingly important role in deep feature representation for visual tasks. Self-attention updates the feature at each position by computing a weighted sum of features using…
The self-attention mechanism, a cornerstone of Transformer-based state-of-the-art deep learning architectures, is largely heuristic-driven and fundamentally challenging to interpret. Establishing a robust theoretical foundation to explain…
Recent studies indicate that hierarchical Vision Transformer with a macro architecture of interleaved non-overlapped window-based self-attention \& shifted-window operation is able to achieve state-of-the-art performance in various visual…
The intimate interaction between phase-matched parametric amplification and strong nonlinear mode coupling in a multimode optical fiber results in a saturable spatial mode conversion that appears to violate the conservation of momentum. We…
Multimode fibers (MMFs) can transmit multiple guided modes simultaneously, making them a promising platform for high-resolution biomedical imaging, endoscopy and high-bandwidth optical communication. However, their complex modal behavior,…
We present lambda layers -- an alternative framework to self-attention -- for capturing long-range interactions between an input and structured contextual information (e.g. a pixel surrounded by other pixels). Lambda layers capture such…
Optical approaches have made great strides towards the goal of high-speed, energy-efficient computing necessary for modern deep learning and AI applications. Read-in and read-out of data, however, limit the overall performance of existing…
The great advances of learning-based approaches in image processing and computer vision are largely based on deeply nested networks that compose linear transfer functions with suitable non-linearities. Interestingly, the most frequently…
Self-attention has emerged as a core component of modern neural architectures, yet its theoretical underpinnings remain elusive. In this paper, we study self-attention through the lens of interacting entities, ranging from agents in…
Self-attention is essential to Transformer architectures, yet how information is embedded in the self-attention matrices and how different objective functions impact this process remains unclear. We present a mathematical framework to…