Related papers: Spelling Rules for the Monster/Semple Tower
The rapid growth of natural language processing (NLP) and pre-trained language models have enabled accurate text classification in a variety of settings. However, text classification models are susceptible to backdoor attacks, where an…
While large-scale neural language models, such as GPT2 and BART, have achieved impressive results on various text generation tasks, they tend to get stuck in undesirable sentence-level loops with maximization-based decoding algorithms…
We investigate the relation between the generating functions of the highest weight vectors of the Monster module and the McKay-Thompson series.
As children enter elementary school, their understanding of the ordinal structure of numbers transitions from a memorized count list of the first 50-100 numbers to knowing the successor function and understanding the countably infinite. We…
We introduce the concept of a type system~$\Part$, that is, a partition on the set of finite words over the alphabet~$\{0,1\}$ compatible with the partial action of Thompson's group~$V$, and associate a subgroup~$\Stab{V}{\Part}$ of~$V$. We…
We address the problem of finding nice labellings for event structures of degree 3. We develop a minimum theory by which we prove that the labelling number of an event structure of degree 3 is bounded by a linear function of the…
The probabilistic interpretation of the standard Regge-Gribov model with triple pomeron interactions is discussed. It is stated that introduction of probabilities within this model is not unique and depends on what is meant under the…
We consider an urn model with multiple drawing and random time-dependent addition matrix. The model is very general with respect to previous literature: the number of sampled balls at each time-step is random, the addition matrix has…
We investigate contributions of spacetime wormholes, describing baby universe emission and absorption, to calculations of entropies and correlation functions, for example those based on the replica method. We find that the rules of the…
Recurrent neural networks are convenient and efficient models for language modeling. However, when applied on the level of characters instead of words, they suffer from several problems. In order to successfully model long-term…
Hypernym Discovery is the task of identifying potential hypernyms for a given term. A hypernym is a more generalized word that is super-ordinate to more specific words. This paper explores several approaches that rely on co-occurrence…
When estimating a proportion and only a sample of triplets is given, dependencies within the triplets are to be accounted for. Without assuming a distribution for the success count of the triplet, together with the proportion, as second and…
Children can use the statistical regularities of their environment to learn word meanings, a mechanism known as cross-situational learning. We take a computational approach to investigate how the information present during each observation…
In this article the idea of random variables over the set theoretic universe is investigated. We explore what it can mean for a random set to have a specific probability of belonging to an antecedently given class of sets.
We consider random binary trees that appear as the output of certain standard algorithms for sorting and searching if the input is random. We introduce the subtree size metric on search trees and show that the resulting metric spaces…
This paper considers the enumeration of ternary trees (i.e. rooted ordered trees in which each vertex has 0 or 3 children) avoiding a contiguous ternary tree pattern. We begin by finding recurrence relations for several simple tree…
The upper bounds that the LHC measurements searching for heavy resonances beyond the Standard model set on the resonance production cross sections are not universal. They depend on various characteristics of the resonance under…
It has recently been demonstrated empirically that in-context learning emerges in transformers when certain distributional properties are present in the training data, but this ability can also diminish upon further training. We provide a…
We study response selection for multi-turn conversation in retrieval-based chatbots. Existing work either concatenates utterances in context or matches a response with a highly abstract context vector finally, which may lose relationships…
Understanding interactions between objects in an image is an important element for generating captions. In this paper, we propose a relationship-based neural baby talk (R-NBT) model to comprehensively investigate several types of pairwise…