Related papers: Temporal Anchoring in Deepening Embedding Spaces: …
The goal of the present work is three-fold. The first goal is to set foundational results on optimal transport in Lorentzian (pre-)length spaces, including cyclical monotonicity, stability of optimal couplings and Kantorovich duality…
Given two independent point processes and a certain rule for matching points between them, what is the fraction of matched points over infinitely long streams? In many application contexts, e.g., secure networking, a meaningful matching…
The purpose of this article is to generalize some known characterizations of Banach space properties in terms of graph preclusion. In particular, it is shown that superreflexivity can be characterized by the non-equi-bi-Lipschitz…
Large pre-trained vision-language models (VLMs), such as CLIP, have shown unprecedented zero-shot performance across a wide range of tasks. Nevertheless, these models may be unreliable under distributional shifts, as their performance is…
For linear parabolic initial-boundary value problems with self-adjoint, time-homogeneous elliptic spatial operator in divergence form with Lipschitz-continuous coefficients, and for incompatible, time-analytic forcing term in…
Post-training large language models (LLMs) often suffers from catastrophic forgetting, where improvements on a target objective degrade previously acquired capabilities. Recent evidence suggests that this phenomenon is primarily driven by…
Adaptivity and local mesh refinement are crucial for the efficient numerical simulation of wave phenomena in complex geometry. Local mesh refinement, however, can impose a tiny time-step across the entire computational domain when using…
Stability and robustness are critical for deploying Transformers in safety-sensitive settings. A principled way to enforce such behavior is to constrain the model's Lipschitz constant. However, approximation-theoretic guarantees for…
Recently (Elkin, Filtser, Neiman 2017) introduced the concept of a {\it terminal embedding} from one metric space $(X,d_X)$ to another $(Y,d_Y)$ with a set of designated terminals $T\subset X$. Such an embedding $f$ is said to have…
Text-to-image diffusion models can generate stunning visuals, yet they often fail at tasks children find trivial--like placing a dog to the right of a teddy bear rather than to the left. When combinations get more unusual--a giraffe above…
A long-standing issue in the parallel-in-time community is the poor convergence of standard iterative parallel-in-time methods for hyperbolic partial differential equations (PDEs), and for advection-dominated PDEs more broadly. Here, a…
We consider the vertex-centered finite volume method with first-order conforming ansatz functions. The adaptive mesh-refinement is driven by the local contributions of the weighted-residual error estimator. We prove that the adaptive…
The paper investigates two inertial extragradient algorithms for seeking a common solution to a variational inequality problem involving a monotone and Lipschitz continuous mapping and a fixed point problem with a demicontractive mapping in…
Timelike entanglement entropy is a complex measure of information that is holographically realized by an appropriate combination of spacelike and timelike extremal surfaces. This measure is highly sensitive to Lorentz invariance breaking.…
We give asymptotically converging semidefinite programming hierarchies of outer bounds on bilinear programs of the form $\mathrm{Tr}\big[M(X\otimes Y)\big]$, maximized with respect to semidefinite constraints on $X$ and $Y$. Applied to the…
Operator learning based on neural operators has emerged as a promising paradigm for the data-driven approximation of operators, mapping between infinite-dimensional Banach spaces. Despite significant empirical progress, our theoretical…
Various forms of sparse attention have been explored to mitigate the quadratic computational and memory cost of the attention mechanism in transformers. We study sparse transformers not through a lens of efficiency but rather in terms of…
We prove that a minimal Transformer with frozen weights emulates a broad class of algorithms by in-context prompting. We formalize two modes of in-context algorithm emulation. In the task-specific mode, for any continuous function $f:…
We consider parallel simulations for asynchronous systems employing L processing elements that are arranged on a ring. Processors communicate only among the nearest neighbors and advance their local simulated time only if it is guaranteed…
We show how the Loschmidt echo of a product state after a quench to a conformal invariant critical point and its leading finite time corrections can be predicted by using conformal field theories (CFT). We check such predictions with tensor…