Related papers: Spelling Rules for the Monster/Semple Tower
Subject-verb agreement in the presence of an attractor noun located between the main noun and the verb elicits complex behavior: judgments of grammaticality are modulated by the grammatical features of the attractor. For example, in the…
The phenomena of in-context learning has typically been thought of as "learning from examples". In this work which focuses on Machine Translation, we present a perspective of in-context learning as the desired generation task maintaining…
We present a modular function-based approach to explaining, for primes larger than 3, the exponents that appear in the prime decomposition of the order of the monster finite simple group.
The wave of pre-training language models has been continuously improving the quality of the machine-generated conversations, however, some of the generated responses still suffer from excessive repetition, sometimes repeating words from…
Long Short-Term Memory recurrent neural network (LSTM) is widely used and known to capture informative long-term syntactic dependencies. However, how such information are reflected in its internal vectors for natural text has not yet been…
We consider how Morita equivalences are compatible with the notion of a corner subring. Namely, we outline a canonical way to replace a corner subring of a given ring with one which is Morita equivalent, and look at how such an equivalence…
Recurrent Neural Networks (RNN), Long Short-Term Memory Networks (LSTM), and Memory Networks which contain memory are popularly used to learn patterns in sequential data. Sequential data has long sequences that hold relationships. RNN can…
Let $I(X,R)$ be the incidence algebra of the preordered set $X$ over the ring $R$. In the case of a finite connected partially ordered set $X$, we prove that the subgroup of inner multiplicative automorphisms is a direct factor of the group…
Logical inference, an integral feature of the Semantic Web, is the process of deriving new triples by applying entailment rules on knowledge bases. The entailment rules are determined by the model-theoretic semantics. Incorporating context…
For a bivariate random vector (X,Y), symmetry conditions are presented that yield stochastic orderings among |X|, |Y|, |max(X,Y)|, and | min(X, Y)|. Partial extensions of these results for multivariate random vectors (X1,...,Xn) are also…
Contextualized representation models such as ELMo (Peters et al., 2018a) and BERT (Devlin et al., 2018) have recently achieved state-of-the-art results on a diverse array of downstream NLP tasks. Building on recent token-level probing work,…
Several problems in machine learning, statistics, and other fields rely on computing eigenvectors. For large scale problems, the computation of these eigenvectors is typically performed via iterative schemes such as subspace iteration or…
In this paper, we consider the iterative method of subspace corrections with random ordering. We prove identities for the expected convergence rate, which can provide sharp estimates for the error reduction per iteration. We also study the…
This paper takes a step towards theoretical analysis of the relationship between word embeddings and context embeddings in models such as word2vec. We start from basic probabilistic assumptions on the nature of word vectors, context…
We begin a systematic study of the enumerative combinatorics of mixed succession rules, which are succession rules such that, in the associated generating tree, the nodes are allowed to produce their sons at several different levels…
This paper deals with the distribution of descent number in standard Young tableaux of certain shapes. A simple explicit formula is presented for the number of tableaux of any shape with two rows, with any specified number of descents. For…
A theoretical analysis is performed of Penning-trap experiments comparing protons and antiprotons to test CPT and Lorentz symmetry through measurements of anomalous magnetic moments and charge-to-mass ratios. Possible CPT and Lorentz…
A common approach for sequence tagging tasks based on contextual word representations is to train a machine learning classifier directly on these embedding vectors. This approach has two shortcomings. First, such methods consider single…
Rich words are characterized by containing the maximum possible number of distinct palindromes. Several characteristic properties of rich words have been studied; yet the analysis of repetitions in rich words still involves some interesting…
We study parameters of the convexity spaces associated with families of sets in $\mathbb{R}^d$ where every intersection between $t$ sets of the family has its Betti numbers bounded from above by a function of $t$. Although the Radon number…