English
Related papers

Related papers: Jordan-RoPE: Non-Semisimple Relative Positional En…

200 papers

The attention module, which is a crucial component in Transformer, cannot scale efficiently to long sequences due to its quadratic complexity. Many works focus on approximating the dot-then-exponentiate softmax function in the original…

Machine Learning · Computer Science 2021-11-04 Shengjie Luo , Shanda Li , Tianle Cai , Di He , Dinglan Peng , Shuxin Zheng , Guolin Ke , Liwei Wang , Tie-Yan Liu

Proximity gaps and correlated agreement have become central tools in the analysis of interactive oracle proofs of proximity (IOPPs) and code-based SNARKs. Informally, a proximity-gap statement says that for a structured set of words -- such…

Information Theory · Computer Science 2026-05-11 Chen Yuan , Ruiqi Zhu

Exceptional points (EPs) are degeneracy of non-Hermitian Hamiltonians, at which the eigenvalues, along with their eigenvectors, coalesce. Their orders are given by the Jordan decomposition. Here, we focus on higher-order EPs arising in…

Quantum Physics · Physics 2023-04-18 Kang Yang , Ipsita Mandal

It was recently suggested -- based on general self-consistency arguments as well as results from the bootstrap (arXiv:2005.07708, arXiv:2007.11539, arXiv:2007.04190) -- that the CFT describing the $Q$-state Potts model is logarithmic for…

Mathematical Physics · Physics 2024-04-01 Lawrence Liu , Jesper Lykke Jacobsen , Hubert Saleur

We study a particular type of logarithmic extension of SL(2,R) Wess-Zumino-Witten models. It is based on the introduction of affine Jordan cells constructed as multiplets of quasi-primary fields organized in indecomposable representations…

High Energy Physics - Theory · Physics 2008-11-26 Jorgen Rasmussen

We complete the determination of the $\ell$-block distribution of characters for quasi-simple exceptional groups of Lie type up to some minor ambiguities relating to non-uniqueness of Jordan decomposition. For this, we first determine the…

Representation Theory · Mathematics 2025-02-13 Radha Kessar , Gunter Malle

Neural language models process sequences of words, but the mathematical operations inside them are insensitive to the order in which words appear. Positional encodings are the component added to remedy this. Despite their importance,…

Machine Learning · Computer Science 2026-04-08 Giansalvo Cirrincione

Approximate Agreement ($\mathcal{AA}$) is a fundamental primitive that, even in the presence of Byzantine faults, allows honest parties to obtain close (but not necessarily identical) outputs that lie within the range of their inputs. While…

Distributed, Parallel, and Cluster Computing · Computer Science 2026-05-07 Marc Fuchs , Diana Ghinea , Zahra Parsaeian , Joel Rybicki

We introduce a novel positional encoding strategy for Transformer-style models, addressing the shortcomings of existing, often ad hoc, approaches. Our framework provides a flexible mapping from the algebraic specification of a domain to an…

Machine Learning · Computer Science 2024-11-01 Konstantinos Kogkalidis , Jean-Philippe Bernardy , Vikas Garg

Position encoding is the primary mechanism which induces notion of sequential order for input tokens in transformer architectures. Even though this formulation in the original transformer paper has yielded plausible performance for general…

Computation and Language · Computer Science 2023-10-10 Eren Unlu

Deep learning-based reduced order models (DL-ROMs) have been recently proposed to overcome common limitations shared by conventional reduced order models (ROMs) - built, e.g., through proper orthogonal decomposition (POD) - when applied to…

Numerical Analysis · Mathematics 2021-11-03 Stefania Fresca , Andrea Manzoni

Long-context large language models (LLMs) have achieved remarkable advancements, driven by techniques like Rotary Position Embedding (RoPE) (Su et al., 2023) and its extensions (Chen et al., 2023; Liu et al., 2024c; Peng et al., 2023). By…

Computation and Language · Computer Science 2025-10-24 Bowen Yang , Bharat Venkitesh , Dwarak Talupuru , Hangyu Lin , David Cairuz , Phil Blunsom , Acyr Locatelli

Relative positional embeddings (RPE) have received considerable attention since RPEs effectively model the relative distance among tokens and enable length extrapolation. We propose KERPLE, a framework that generalizes relative position…

Computation and Language · Computer Science 2022-10-14 Ta-Chung Chi , Ting-Han Fan , Peter J. Ramadge , Alexander I. Rudnicky

Without positional information, attention-based Transformer neural networks are permutation-invariant. Absolute or relative positional embeddings are the most popular ways to feed Transformer models with positional information. Absolute…

Machine Learning · Computer Science 2021-11-10 Tatiana Likhomanenko , Qiantong Xu , Gabriel Synnaeve , Ronan Collobert , Alex Rogozhnikov

Currently, many studies view DNA sequences as a special type of language and utilize Transformers to model them. These studies use fixed-length k-mer segmentation and BPE subword tokenization but lack a systematic evaluation to determine…

Computation and Language · Computer Science 2025-07-22 Chenlei Gong , Yuanhe Tian , Lei Mao , Yan Song

Positional embeddings (PE) play a crucial role in Vision Transformers (ViTs) by providing spatial information otherwise lost due to the permutation invariant nature of self attention. While absolute positional embeddings (APE) have shown…

Computer Vision and Pattern Recognition · Computer Science 2026-04-14 Md Abtahi Majeed Chowdhury , Md Rifat Ur Rahman , Akil Ahmad Taki

Loop closures are essential for correcting odometry drift and creating consistent maps, especially in the context of large-scale navigation. Current methods using dense point clouds for accurate place recognition do not scale well due to…

Computer Vision and Pattern Recognition · Computer Science 2025-07-29 Débora N. P. Oliveira , Joshua Knights , Sebastián Barbas Laina , Simon Boche , Wolfram Burgard , Stefan Leutenegger

The theories of $\pi$-points and modules of constant Jordan type have been a topic of much recent interest in the field of finite group scheme representation theory. These theories allow for a finite group scheme module $M$ to be restricted…

Representation Theory · Mathematics 2015-09-07 Andrew J. Talian

We present an algorithm to compute the Jordan chain of a nearly defective matrix with a $2\times2$ Jordan block. The algorithm is based on an inverse-iteration procedure and only needs information about the invariant subspace corresponding…

Numerical Analysis · Mathematics 2017-04-25 Felipe Hernández , Adi Pick , Steven G. Johnson

Traditional reduced order modeling techniques such as the reduced basis (RB) method (relying, e.g., on proper orthogonal decomposition (POD)) suffer from severe limitations when dealing with nonlinear time-dependent parametrized PDEs,…

Numerical Analysis · Mathematics 2020-01-14 Stefania Fresca , Luca Dede , Andrea Manzoni
‹ Prev 1 3 4 5 6 7 10 Next ›