Related papers: Topological Data Analysis in Information Space
Ordinal data analysis is an interesting direction in machine learning. It mainly deals with data for which only the relationships `$<$', `$=$', `$>$' between pairs of points are known. We do an attempt of formalizing structures behind…
We build information geometry for a partially ordered set of variables and define the orthogonal decomposition of information theoretic quantities. The natural connection between information geometry and order theory leads to efficient…
We investigate the pertinence of methods from algebraic topology for text data analysis. These methods enable the development of mathematically-principled isometric-invariant mappings from a set of vectors to a document embedding, which is…
Inferring topological and geometrical information from data can offer an alternative perspective on machine learning problems. Methods from topological data analysis, e.g., persistent homology, enable us to obtain such information,…
Topological measurements are increasingly being accepted as an important tool for quantifying complex structures. In many applications, these structures can be expressed as nodal domains of real-valued functions and are obtained only…
In this paper we develop some combinatorial models for continuous spaces. In this spirit we study the approximations of continuous spaces by graphs, molecular spaces and coordinate matrices. We define the dimension on a discrete space by…
The incredible variety of galaxy shapes cannot be summarized by human defined discrete classes of shapes without causing a possibly large loss of information. Dictionary learning and sparse coding allow us to reduce the high dimensional…
Any symmetric affinity function $w: V\times V \to \mathbb{R}_+$ defined on a discrete set $V$ induces Euclidean space structure on $V$. In particular, an undirected graph specified by an affinity (or adjacency) matrix can be considered as a…
We develop and test a novel unsupervised algorithm for word sense induction and disambiguation which uses topological data analysis. Typical approaches to the problem involve clustering, based on simple low level features of distance in…
The manifold of empirical mean values of statistical data ad infinitum has a geometric shape that depends on the probability measure that governs the generating model. Large deviation theory produces entropy functions that depend on both…
Improvements in computational and experimental capabilities are rapidly increasing the amount of scientific data that is routinely generated. In applications that are constrained by memory and computational intensity, excessively large…
How does the topological space of science emerge? Inspired by the concept of maps of science, i.e. mapping scientific topics to a scientific space, we ask which topological structure a dynamical process of authors collaborating and…
Information distance is a parameter-free similarity measure based on compression, used in pattern recognition, data mining, phylogeny, clustering, and classification. The notion of information distance is extended from pairs to multiples…
Topological data analysis (TDA) is a rising branch in modern applied mathematics. It extracts topological structures as features of a given space and uses these features to analyze digital data. Persistent homology, one of the central tools…
Observations on the past provide some hints about what will happen in the future, and this can be quantified using information theory. The ``predictive information'' defined in this way has connections to measures of complexity that have…
Density map is an effective visualization technique for depicting the scalar field distribution in 2D space. Conventional methods for constructing density maps are mainly based on Euclidean distance, limiting their applicability in urban…
Features such as photon rings, jets, or hot. spots can leave particular topological signatures in a black hole image. As such, topological data analysis can be used to characterize images resulting from high resolution observations…
In many problems in data mining and machine learning, data items that need to be clustered or classified are not points in a high-dimensional space, but are distributions (points on a high dimensional simplex). For distributions, natural…
Techniques from computational topology, in particular persistent homology, are becoming increasingly relevant for data analysis. Their stable metrics permit the use of many distance-based data analysis methods, such as multidimensional…
Complex networks are characterized by latent geometries induced by their topology or by the dynamics on the top of them. In the latter case, different network-driven processes induce distinct geometric features that can be captured by…