Related papers: Spelling Rules for the Monster/Semple Tower
In this article we will apply complex projective metrics to sequences of complex transfer operators generated by Young towers, countable shifts and other types of distance expanding maps (possibly time dependent) with countable degrees. We…
Naming game simulates the process of naming an object by a single word, in which a population of communicating agents can reach global consensus asymptotically through iteratively pair-wise conversations. We propose an extension of the…
We introduce a new method for studying murmurations, based on random matrix theory. With this method, we exhibit murmurations or similar phenomena: assuming ratios conjectures, for elliptic curves ordered by height, quadratic twists of a…
Many structured prediction problems (particularly in vision and language domains) are ambiguous, with multiple outputs being correct for an input - e.g. there are many ways of describing an image, multiple ways of translating a sentence;…
To ensure large language models (LLMs) are used safely, one must reduce their propensity to hallucinate or to generate unacceptable answers. A simple and often used strategy is to first let the LLM generate multiple hypotheses and then…
Recursive learning -- where models are trained on data generated by previous versions of themselves -- is increasingly common in large language models, autonomous agents, and self-supervised systems. However, standard performance metrics…
The traditional way of sentence-level event detection involves two important subtasks: trigger identification and trigger classifications, where the identified event trigger words are used to classify event types from sentences. However,…
Small and mid-sized generative language models have gained increasing attention. Their size and availability make them amenable to being analyzed at a behavioral as well as a representational level, allowing investigations of how these…
We study the restricted families of orthogonal projections in $\mathbb{R}^3$. We show that there are families of random subspaces which admit a Marstrand- Mattila type projection theorem.
Recurrent Neural Networks (RNNs) have been shown to capture various aspects of syntax from raw linguistic input. In most previous experiments, however, learning happens over unrealistic corpora, which do not reflect the type and amount of…
We focus on working on incidence rings, a class of (possibly infinite) matrix rings indexed by ordered sets. Some general properties about them are given, including how they are always the inverse limit of finite matrix rings, giving a…
We classify all finite groups such that all irreducible character degrees appear with multiplicity at most $2$. As a consequence, we prove that the largest group with at most $2$ irreducible characters of the same degree is the Baby…
Small bodies in the solar system are conventionally classified into asteroids and comets. However, it is recently found that a small number of objects can exhibit properties of both asteroids and comets. Some are more consistent with…
We study word learning in subword and character language models with the psycholinguistic lexical decision task. While subword LMs struggle to discern words and non-words with high accuracy, character LMs solve this task easily and…
We study the class of word-building games, where two players pick letters from a finite alphabet to construct a finite or infinite word. The outcome is determined by whether the resulting word lies in a prescribed set (a win for player $A$)…
In this paper, we study three matching problems all of which came up quite recently in the field of machine teaching. The cost of a matching is defined in such a way that, for some formal model of teaching, it equals (or bounds) the number…
We identify the similarity between two words in English by casting the task as machine translation performance prediction (MTPP) between the words given the context and the distance between their similarities. We use referential translation…
Scaling laws aim to accurately predict model performance across different scales. Existing scaling-law studies almost exclusively rely on cross-entropy as the evaluation metric. However, cross-entropy provides only a partial view of…
A monster is an automaton in which every function from states to states is represented by at least one letter. A modifier is a set of functions allowing one to transform a set of automata into one automaton. We revisit some language…
Recently, several methods have been proposed to explain the predictions of recurrent neural networks (RNNs), in particular of LSTMs. The goal of these methods is to understand the network's decisions by assigning to each input variable,…