Related papers: Jordan-RoPE: Non-Semisimple Relative Positional En…
POD-DL-ROMs have been recently proposed as an extremely versatile strategy to build accurate and reliable reduced order models (ROMs) for nonlinear parametrized partial differential equations, combining (i) a preliminary dimensionality…
Positional encoding (PE) underpins how permutation-invariant Transformers represent sequence order, yet how positional information is processed and stored remains poorly understood. Modern PE methods such as RoPE still struggle on tasks…
Let $G$ be a symplectic group over a nonarchimedean local field of characteristic zero and odd residual characteristic. Given an irreducible cuspidal representation of G, we determine its Langlands parameter (equivalently, its Jordan blocks…
3D visual grounding aims to identify objects in 3D point cloud scenes that match specific natural language descriptions. This requires the model to not only focus on the target object itself but also to consider the surrounding environment…
In this paper, we study the realizability problem for retarded functional differential equations near an equilibrium point undergoing a nonlinear mode interaction between a saddle-node bifurcation and a non-resonant multiple Hopf…
The attention mechanism is a core primitive in modern large language models (LLMs) and AI more broadly. Since attention by itself is permutation-invariant, position encoding is essential for modeling structured domains such as language.…
The efficacy of Multimodal Transformers in visually-rich document understanding (VrDU) is critically constrained by two inherent limitations: the lack of explicit modeling for logical reading order and the interference of visual tokens that…
A special class of Jordan algebras over a field $F$ of characteristic zero is considered. Such an algebra consists of an $r$-dimensional subspace of the vector space of all square matrices of a fixed order $n$ over $F$. It contains the…
Rotary Position Embedding (RoPE) is an efficient position encoding approach and is widely utilized in numerous large language models (LLMs). Recently, a lot of methods have been put forward to further expand the context window based on…
This paper studies how Transformer models with Rotary Position Embeddings (RoPE) develop emergent, wavelet-like properties that compensate for the positional encoding's theoretical limitations. Through an analysis spanning model scales,…
For revDSD double hybrids, the G\"orling-Levy second-order perturbation theory component is an Achilles' Heel when applied to systems with significant near-degeneracy ("static") correlation. We have explored its replacement by the direct…
Let G be a simple simple-connected exceptional algebraic group of type G_2, F_4, E_6 or E_7 over an algebraically closed field k of characteristic p>0 with \g=Lie(G). For each nilpotent orbit G.e of \g, we list the Jordan blocks of the…
Rotary Position Embeddings (RoPE) have been shown to effectively encode positional information in transformer-based language models. However, these models fail to generalize past the sequence length they were trained on. We present YaRN…
The synchronization problem over the special orthogonal group $SO(d)$ consists of estimating a set of unknown rotations $R_1,R_2,...,R_n$ from noisy measurements of a subset of their pairwise ratios $R_{i}^{-1}R_{j}$. The problem has found…
Although large language models (LLMs) have achieved significant progress in handling long-context inputs, they still suffer from the ``lost-in-the-middle'' problem, where crucial information in the middle of the context is often…
Tensor Attention extends traditional attention mechanisms by capturing high-order correlations across multiple modalities, addressing the limitations of classical matrix-based attention. Meanwhile, Rotary Position Embedding…
We present a new class of efficient attention mechanisms applying universal 3D Relative Positional Encoding (RPE) methods given by arbitrary integrable modulation functions $f$. They lead to the new class of 3D-Transformer models, called…
In this article we prove that the elliptic, hyperbolic and nilpotent (or unipotent) additive (or multiplicative) Jordan components of an endomorphism $X$ (or an isomorphism $g$) of a finite dimensional vector space are given by polynomials…
Relative positional encoding is widely used in vanilla and linear transformers to represent positional information. However, existing encoding methods of a vanilla transformer are not always directly applicable to a linear transformer,…
One essential ingredient in many machine learning (ML) based methods for atomistic modeling of materials and molecules is the use of locality. While allowing better system-size scaling, this systematically neglects long-range (LR) effects,…