Related papers: Complete spelling rules for the Monster tower over…
Phrase structure trees have a hierarchical structure. In many subjects, most notably in Taxonomy such tree structures have been studied using ultrametrics. Here syntactical hierarchical phrase trees are subject to a similar analysis, which…
A big convergence of model architectures across language, vision, speech, and multimodal is emerging. However, under the same name "Transformers", the above areas use different implementations for better performance, e.g., Post-LayerNorm…
An important question concerning contextualized word embedding (CWE) models like BERT is how well they can represent different word senses, especially those in the long tail of uncommon senses. Rather than build a WSD system as in previous…
Designers of statistical machine translation (SMT) systems have begun to employ tree-structured translation models. Systems involving tree-structured translation models tend to be complex. This article aims to reduce the conceptual…
Naming game simulates the process of naming an object by a single word, in which a population of communicating agents can reach global consensus asymptotically through iteratively pair-wise conversations. We propose an extension of the…
A real Bott manifold is the total space of a sequence of $\R P^1$ bundles starting with a point, where each $\R P^1$ bundle is projectivization of a Whitney sum of two real line bundles. A real Bott manifold is a real toric manifold which…
A support vector machine (SVM) is an algorithm that finds a hyperplane which optimally separates labeled data points in $\mathbb{R}^n$ into positive and negative classes. The data points on the margin of this separating hyperplane are…
We investigate the geometry of predictive information across the layers of large language models (LLMs). We repurpose representation lenses-learned affine maps trained to predict the next token from intermediate residual streams-as…
Assuming Hartshorne's conjecture on complete intersections, we classify projective bundles over projective spaces which has a smooth blow up structure over another projective space. Under some assumptions, we also classify projective…
We carefully present an elementary proof of the well known theorem that each homotopy group (or, in degree zero, pointed set) of the inverse limit of a tower of fibrations maps naturally onto the inverse limit of the homotopy groups (or, in…
Recent studies on neural networks with pre-trained weights (i.e., BERT) have mainly focused on a low-dimensional subspace, where the embedding vectors computed from input words (or their contexts) are located. In this work, we propose a new…
Consider a weighted branching process generated by a point process on $[0,1]$, whose atoms sum up to one. Then the weights of all individuals in any given generation sum up to one, as well. We define a nested occupancy scheme in random…
In neural machine translation, a source sequence of words is encoded into a vector from which a target sequence is generated in the decoding phase. Differently from statistical machine translation, the associations between source words and…
While machine translation (MT) systems are achieving increasingly strong performance on benchmarks, they often produce translations with errors and anomalies. Understanding these errors can potentially help improve the translation quality…
We construct most symmetric Saddle towers in Heisenberg space i.e. periodic minimal surfaces that can be seen as the desingularization of vertical planes intersecting equiangularly. The key point is the construction of a suitable barrier to…
The Clifford hierarchy is a fundamental structure in quantum computation whose mathematical properties are not fully understood. In this work, we characterize permutation gates -- unitaries which permute the $2^n$ basis states -- in the…
A mixed Steiner system MS$(t,k,Q)$ is a set (code) $C$ of words of weight $k$ over an alphabet $Q$, where not all coordinates of a word have the same alphabet size, each word of weight $t$, over $Q$, has distance $k-t$ from exactly one…
This is an experiential study of investigating a consistent method for deriving the correlation between sentence vector and semantic meaning of a sentence. We first used three state-of-the-art word/sentence embedding methods including…
Automata-logic connections are pillars of the theory of regular languages. Such connections are harder to obtain for transducers, but important results have been obtained recently for word-to-word transformations, showing that the three…
Cross-lingual word vectors are typically obtained by fitting an orthogonal matrix that maps the entries of a bilingual dictionary from a source to a target vector space. Word vectors, however, are most commonly used for sentence or…