Related papers: Lossless Prioritized Embeddings
While raw cosine similarity in pretrained embedding spaces exhibits strong rank correlation with human judgments, anisotropy induces systematic miscalibration of absolute values: scores concentrate in a narrow high-similarity band…
We first prove that for every metrizable space $X$, for every closed subset $F$ whose complement is zero-dimensional, the space $X$ can be embedded into a product space of the closed subset $F$ and a metrizable zero-dimensional space as a…
Since persistence diagrams do not admit an inner product structure, a map into a Hilbert space is needed in order to use kernel methods. It is natural to ask if such maps necessarily distort the metric on persistence diagrams. We show that…
The optimal Orlicz target space is exhibited for embeddings of fractional-order Orlicz-Sobolev spaces in $\mathbb R^n$. An improved embedding with an Orlicz-Lorentz target space, which is optimal in the broader class of all…
Embedded spaces are a key feature in deep learning. Good embedded spaces represent the data well to support classification and advanced techniques such as open-set recognition, few-short learning and explainability. This paper presents a…
In applications with significant class imbalance or asymmetric costs, metrics such as the $F_\beta$-measure, AM measure, Jaccard similarity coefficient, and weighted accuracy offer more suitable evaluation criteria than standard binary…
Embedding complex objects as vectors in low dimensional spaces is a longstanding problem in machine learning. We propose in this work an extension of that approach, which consists in embedding objects as elliptical probability…
Modern AI is opening the door to collective decision-making in which participants express their views as free-form text rather than voting on a fixed set of candidates. A natural idea is to embed these opinions in a vector space so that the…
Typical graph embeddings may not capture type-specific bipartite graph features that arise in such areas as recommender systems, data visualization, and drug discovery. Machine learning methods utilized in these applications would be better…
We show that, for a separable and complete metric space $M$, the Lipschitz-free space $\mathcal F(M)$ embeds linearly and almost-isometrically into $\ell_1$ if and only if $M$ is a subset of an $\mathbb R$-tree with length measure 0.…
Let $G$ be an Orlicz function and let $ \alpha, \beta, s$ be positive real numbers. Under certain conditions on the Orlicz function $ G $, we establish some continuous embeddings results between the fractional order Orlicz-Sobolev spaces…
This work studies an explicit embedding of the set of probability measures into a Hilbert space, defined using optimal transport maps from a reference probability density. This embedding linearizes to some extent the 2-Wasserstein space,…
Metric spaces $(X, d)$ are ubiquitous objects in mathematics and computer science that allow for capturing (pairwise) distance relationships $d(x, y)$ between points $x, y \in X$. Because of this, it is natural to ask what useful…
This paper is concerned with proving some embeddings of the form \begin{equation*} F_{p_{1},q}^{s_{1}}\cdot B_{p_{2},\infty }^{s_{2}}\cdot ...\cdot B_{p_{m},\infty }^{s_{m}}\hookrightarrow F_{p,q}^{s_{1}},\quad m\geq 2. \end{equation*} The…
We study first-order optimization algorithms under the constraint that the descent direction is quantized using a pre-specified budget of $R$-bits per dimension, where $R \in (0 ,\infty)$. We propose computationally efficient optimization…
We give sufficient conditions for a metric space to bilipschitz embed in L_1. In particular, if X is a length space and there is a Lipschitz map u:X--->R such that for every interval I in R, the connected components of the inverse image…
We show that for every $\alpha > 0$, there exist $n$-point metric spaces (X,d) where every "scale" admits a Euclidean embedding with distortion at most $\alpha$, but the whole space requires distortion at least $\Omega(\sqrt{\alpha \log…
Unsupervised representation learning methods are widely used for gaining insight into high-dimensional, unstructured, or structured data. In some cases, users may have prior topological knowledge about the data, such as a known cluster…
Let $(X,\left\Vert \cdot \right\Vert )$ be a real normed space of dimension $N\in \mathbb{N}$ with a basis $(e_{i})_{1}^{N}$ such that the norm is invariant under coordinate permutations. Assume for simplicity that the basis constant is at…
This paper studies the minimal dimension required to embed subset memberships ($m$ elements and ${m\choose k}$ subsets of at most $k$ elements) into vector spaces, denoted as Minimal Embeddable Dimension (MED). The tight bounds of MED are…