Related papers: A linear transformation to accelerate the converge…
The dominant approach to sequence generation is to produce a sequence in some predefined order, e.g. left to right. In contrast, we propose a more general model that can generate the output sequence by inserting tokens in any arbitrary…
Transformers excel at discovering patterns in sequential data, yet their fundamental limitations and learning mechanisms remain crucial topics of investigation. In this paper, we study the ability of Transformers to learn pseudo-random…
We give an algorithm to compute the series expansion for the inverse of a given function. The algorithm is extremely easy to implement and gives the first $N$ terms of the series. We show several examples of its application in calculating…
Transformers have the capacity to act as supervised learning algorithms: by properly encoding a set of labeled training ("in-context") examples and an unlabeled test example into an input sequence of vectors of the same dimension, the…
We define a mapping from transition-based parsing algorithms that read sentences from left to right to sequence labeling encodings of syntactic trees. This not only establishes a theoretical relation between transition-based parsing and…
We observe that the characteristic polynomial of a linearly perturbed semidefinite matrix can be used to give the convergence rate of alternating projections for the positive semidefinite cone and a line. As a consequence, we show that such…
When a sequence of numbers is slowly converging, it can be transformed into a new sequence which, under some assumptions, could converge faster to the same limit. One of the most well--known sequence transformation is Shanks transformation…
This paper is a study of power series, where the coefficients are binomial expressions (iterated finite differences). Our results can be used for series summation, for series transformation, or for asymptotic expansions involving Stirling…
A method is developed for calculating effective sums of divergent series. This approach is a variant of the self-similar approximation theory. The novelty here is in using an algebraic transformation with a power providing the maximal…
The increasing demand for Fourier transforms on geometric algebras has resulted in a large variety. Here we introduce one single straight forward definition of a general geometric Fourier transform covering most versions in the literature.…
The beta integral is applied to accelerate the hypergeometric function $2 F 1\left\{1, B; C ; w\right\}$ to derive new infinite series for constants such as $\pi$ and values of the gamma function. A compendium of new infinite series is…
The purpose of this paper is to discuss the construction of a linear operator, referred to as the bubble transform, which maps scalar functions defined on a bounded domain $\Omega$ in $\mathbb{R}^n$ into a collection of functions with local…
The double-direction orthogonalization algorithm is applied to construct sequences of polynomials, which are orthogonal over the interval [0,1]with the weighting function 1. Functional and recurrent relations are derived for the sequences…
Let {X(t)} be a stationary time series with a.e. positive spectrum. Two consequences of that the bispectrum of {X(t)} is real-valued but nonzero: 1) if {X(t)} is also linear, then it is reversible; 2) {X(t),} can not be causal linear. A…
We propose a framework for sequence-to-sequence contrastive learning (SeqCLR) of visual representations, which we apply to text recognition. To account for the sequence-to-sequence structure, each feature map is divided into different…
The Lie linearizability criteria are extended to complex functions for complex ordinary differential equations. The linearizability of complex ordinary differential equations is used to study the linearizability of corresponding systems of…
A number of problems in the processing of sound and natural language, as well as in other areas, can be reduced to simultaneously reading an input sequence and writing an output sequence of generally different length. There are well…
Linear Transformers and State Space Models have emerged as efficient alternatives to softmax Transformers for causal sequence modeling, enabling parallel training via matrix multiplication and efficient RNN-style inference. However, despite…
This paper is devoted to the study of the log-convexity of combinatorial sequences. We show that the log-convexity is preserved under componentwise sum, under binomial convolution, and by the linear transformations given by the matrices of…
Using linear projections one gets new inequalities for the successive minima of the lattice of sections of an hermitian line bundle on an arithmetic surface.