Related papers: TokenBlowUp: Resolving Representational Singularit…
Many geometry processing pipelines implicitly assume their input data is a manifold, or is sampled from one, with a unique tangent plane at every point. Geometric data, however, routinely contains sharp features like edges, corners,…
Geometric treatments of blow-up solutions for autonomous ordinary differential equations and their blow-up rates are concerned. Our approach focuses on the type of invariant sets at infinity via compactifications of phase spaces, and…
The geometric evolution of token representations in large language models (LLMs) presents a fundamental paradox: while human language inherently organizes semantic information in low-dimensional spaces ($\sim 10^1$ dimensions), modern LLMs…
Large Language Models (LLMs) perform internal computations in continuous vector spaces yet produce discrete tokens -- a fundamental mismatch whose geometric consequences remain poorly understood. We develop a mathematical framework that…
We first introduce and study the notion of multi-weighted blow-ups, which is later used to systematically construct an explicit yet efficient algorithm for functorial logarithmic resolution in characteristic zero, in the sense of Hironaka.…
Stack-theoretic blow-ups have proven to be efficient in resolving singularities over fields of characteristic zero. In this article, we move forward towards positive characteristic where new challenges arise. In particular, the dimension of…
A central challenge in developing Multimodal Large Language Models (MLLMs) is effectively integrating heterogeneous inputs into a cohesive reasoning engine. Current paradigms predominantly rely on modular architectures that introduce…
Real blow-up, including inhomogeneous versions, of boundary faces of a manifold (with corners) is an important tool for resolving singularities, degeneracies and competing notions of homogeneity. These constructions are shown to be…
Large language models (LLMs) reason over discrete token ID sequences, yet modern subword tokenizers routinely produce non-unique encodings: multiple token ID sequences can detokenize to identical surface strings. This representational…
Large Language Models (LLMs) drive current AI breakthroughs despite very little being known about their internal representations. In this work, we propose to shed the light on LLMs inner mechanisms through the lens of geometry. In…
The manifold hypothesis, which assumes that data lies on or close to an unknown manifold of low intrinsic dimension, is a staple of modern machine learning research. However, recent work has shown that real-world data exhibits distinct…
Discrete motion tokenization has recently enabled Large Language Models (LLMs) to serve as versatile backbones for motion understanding and motion-language reasoning. However, existing pipelines typically decouple motion quantization from…
Heterotic orbifold models are promising candidates for models with MSSM like spectra. But orbifolds only correspond to a special place in moduli space, the bigger picture is described by the moduli space of Calabi-Yau spaces. In this talk…
Geometric singular perturbation theory provides a powerful mathematical framework for the analysis of 'stationary' multiple time-scale systems which possess a critical manifold, i.e. a smooth manifold of steady states for the limiting fast…
Blowing up a point p in a manifold M builds a new manifold M' in which p is replaced by the projectivization of the tangent space of M at p. This well-known operation also applies to fixed points of diffeomorphisms, yielding continuous…
Given a singular hypersurface in a regular 2-dimensional scheme essentially of finite type over a field, we construct an embedded resolution of singularities by weighted blow-ups. This differs from our earlier work which required…
We prove the existence of resolution of singularities for arbitrary (not necessarily reduced or irreducible) excellent two-dimensional schemes, via permissible blow-ups. The resolution is canonical, and functorial with respect to…
A full understanding of the behavior of a large language model (LLM) requires our grasp of its input token space. If this space differs from our assumptions, our comprehension of and conclusions about the LLM will likely be flawed. We…
Tokenisation is the first step in almost all NLP tasks, and state-of-the-art transformer-based language models all use subword tokenisation algorithms to process input text. Existing algorithms have problems, often producing tokenisations…
This work introduces topological regularization as a framework for handling ultraviolet divergences in quantum field theory, reinterpreting infinities as topological obstructions at spacetime boundaries. Through geometric compactification…