Related papers: Complete spelling rules for the Monster tower over…
Neural machine translation (NMT) systems are usually trained on a large amount of bilingual sentence pairs and translate one sentence at a time, ignoring inter-sentence information. This may make the translation of a sentence ambiguous or…
A lower bound on the amount of noise that must be added to a GHZ-like entangled state to make it separable (also called the random robustness) is found using the transposition condition. The bound is applicable to arbitrary numbers of…
This paper begins with a description of methods for estimating image probability density functions that reflects the observation that such data is usually constrained to lie in restricted regions of the high-dimensional image space-not…
We study projectivity of moduli spaces on the DT/PT wall crossing in Bridgeland and polynomial stability on a smooth, projective threefold. First, we construct a globally generated line bundle on the moduli stack of higher-rank…
The purpose of this paper is to put the description of number scaling and its effects on physics and geometry on a firmer foundation, and to make it more understandable. A main point is that two different concepts, number and number value…
We study the geometry of Bott towers in the context of toric geometry, describing their associated fans arising from crosspolytopes. We compute the cohomology ring of each stage of the tower, and provide all monomial identities defining…
We provide a new set of on-shell recursion relations for tree-level scattering amplitudes, which are valid for any non-trivial theory of massless particles. In particular, we reconstruct the scattering amplitudes from (a subset of) their…
A Bott manifold is the total space of some iterated $\mathbb C P^1$-bundle over a point. We prove that any graded ring isomorphism between the cohomology rings of two Bott manifolds preserves their Pontrjagin classes. Moreover, we prove…
Two-Tower Vision-Language (VL) models have shown promising improvements on various downstream VL tasks. Although the most advanced work improves performance by building bridges between encoders, it suffers from ineffective layer-by-layer…
Knowledge-based machine translation (KBMT) systems have achieved excellent results in constrained domains, but have not yet scaled up to newspaper text. The reason is that knowledge resources (lexicons, grammar rules, world models) must be…
In this paper, we construct a stratification tower for the equivariant slice filtration. This tower stratifies the slice spectral sequence of a $G$-spectrum $X$ into distinct regions. Within each of these regions, the differentials are…
Sentence embedding is an important research topic in natural language processing (NLP) since it can transfer knowledge to downstream tasks. Meanwhile, a contextualized word representation, called BERT, achieves the state-of-the-art…
A class of self-similar sets of entangled quantum states is introduced, for which a recursive definition is provided. These sets, the "Bell gems," are defined by the subsystem exchange symmetry characteristic of the Bell states. Each Bell…
A great proportion of sequence-to-sequence (Seq2Seq) models for Neural Machine Translation (NMT) adopt Recurrent Neural Network (RNN) to generate translation word by word following a sequential order. As the studies of linguistics have…
The main result is a wall crossing formula for central projections defined on submanifolds of a real projective space. Our formula gives the jump of the degree of such a projection when the center of the projection varies. The fact that the…
While multi-modal large language models (MLLMs) have shown significant progress on many popular visual reasoning benchmarks, whether they possess abstract visual reasoning abilities remains an open question. Similar to the Sudoku puzzles,…
We construct vector bundles $R^r_\mu$ on a smooth projective curve $X$ having the property that for all sheaves $E$ of slope $\mu$ and rank $r$ on $X$ we have an equivalence: $E$ is a semistable vector bundle $\iff$ $Hom(R^r_\mu,E)=0$. As a…
Word and phrase tables are key inputs to machine translations, but costly to produce. New unsupervised learning methods represent words and phrases in a high-dimensional vector space, and these monolingual embeddings have been shown to…
We propose a new method for projective dependency parsing based on headed spans. In a projective dependency tree, the largest subtree rooted at each word covers a contiguous sequence (i.e., a span) in the surface order. We call such a span…
When trained on language data, do transformers learn some arbitrary computation that utilizes the full capacity of the architecture or do they learn a simpler, tree-like computation, hypothesized to underlie compositional meaning systems…