Related papers: Space filling positionality and the Spiroformer
The study of fundamental optics effects has been stimulated through the increasing ability to structure light in all its degrees of freedom (DOFs) in sophisticated but simple experimental settings. However, with such an increase in…
Despite their central role in the success of foundational models and large-scale language modeling, the theoretical foundations governing the operation of Transformers remain only partially understood. Contemporary research has largely…
Following arXiv:1501.03019 [hep-th], we study de Sitter space and spherical subregions on a constant boundary Euclidean time slice of the future boundary in the Poincare slicing. We show that as in that case, complex extremal surfaces exist…
Graph Transformers excel in long-range dependency modeling, but generally require quadratic memory complexity in the number of nodes in an input graph, and hence have trouble scaling to large graphs. Sparse attention variants such as…
We present an exact, analytic solution of the spin dependent quantum transport problem with spin-orbit interaction in a one-dimensional mesoscopic ring with one input and two output leads. We demonstrate that for appropriate parameters…
Transformers have demonstrated remarkable success in sequence modeling, yet effectively incorporating positional information remains a challenging and active area of research. In this paper, we introduce JoFormer, a journey-based…
In previous work, a class of noninvertible topological dynamical systems $f: X \to X$ was introduced and studied; we called these {\em topologically coarse expanding conformal} systems. To such a system is naturally associated a preferred…
Transformers have become the de-facto standard model in artificial intelligence since 2017 despite numerous shortcomings ranging from energy inefficiency to hallucinations. Research has made a lot of progress in improving elements of…
As an inverse problem, we recover the topology of the effective spacetime that a system lies in, in an operational way. This means that from a series of experiments we get a set of points corresponding to events. This continues the previous…
Many machine learning tasks such as multiple instance learning, 3D shape recognition, and few-shot image classification are defined on sets of instances. Since solutions to such problems do not depend on the order of elements of the set,…
In this paper, we describe in detail a model of geometric-functional variability between fshapes. These objects were introduced for the first time by the authors in [Charlier et al. 2015] and are basically the combination of classical…
Accurate and physically consistent modeling of Earth system dynamics requires machine-learning architectures that operate directly on continuous geophysical fields and preserve their underlying geometric structure. Here we introduce…
Two dimensional space-filling bearings are dense packings of disks that can rotate without slip. We consider the entire first family of bearings for loops of size four and propose a hierarchical construction of their contact network. We…
Recently, deep-learning-based approaches have been widely studied for deformable image registration task. However, most efforts directly map the composite image representation to spatial transformation through the convolutional neural…
The Transformer and its variants have been proven to be efficient sequence learners in many different domains. Despite their staggering success, a critical issue has been the enormous number of parameters that must be trained (ranging from…
As is well-known, given the complex sphere P^1 minus two points, there exist nonconstant holomorphic maps from the plane into this set, the simplest example of which is given by applying the exponential map and then composing with a…
We show that a constant number of self-attention layers can efficiently simulate, and be simulated by, a constant number of communication rounds of Massively Parallel Computation. As a consequence, we show that logarithmic depth is…
In this study, we introduce and explore a delay differential equation that lends itself to explicit solutions in the Fourier-transformed space. Through the careful alignment of the initial function, we can construct a highly accurate…
Large Transformer models yield impressive results on many tasks, but are expensive to train, or even fine-tune, and so slow at decoding that their use and study becomes out of reach. We address this problem by leveraging sparsity. We study…
Many deep neural network architectures loosely based on brain networks have recently been shown to replicate neural firing patterns observed in the brain. One of the most exciting and promising novel architectures, the Transformer neural…