English
Related papers

Related papers: Visualizing departures from marginal homogeneity f…

200 papers

This work is concerned with the continuum limit of a graph-based data visualization technique called the t-Distributed Stochastic Neighbor Embedding (t-SNE), which is widely used for visualizing data in a variety of applications, but is…

Machine Learning · Statistics 2026-04-15 Jeff Calder , Zhonggan Huang , Ryan Murray , Adam Pickarski

Detection of material inhomogeneities is an important task in magnetic imaging and plays a significant role in understanding physical processes. For example, in spintronics, the sample heterogeneity determines the onset of current-driven…

Computational Physics · Physics 2021-11-03 Illia Horenko , Davi Rodrigues , Terence O'Kane , Karin Everschor-Sitte

The bottom-up saliency, an early stage of humans' visual attention, can be considered as a binary classification problem between centre and surround classes. Discriminant power of features for the classification is measured as mutual…

Computer Vision and Pattern Recognition · Computer Science 2013-06-07 Anh Cat Le Ngo , Kenneth Li-Minn Ang , Guoping Qiu , Jasmine Kah-Phooi Seng

Trajectory Inference (TI) seeks to recover latent dynamical processes from snapshot data, where only independent samples from time-indexed marginals are observed. In applications such as single-cell genomics, destructive measurements make…

Machine Learning · Computer Science 2026-04-23 Chao Wang , Luca Nepote , Giulio Franzese , Pietro Michiardi

The bottom-up saliency, an early stage of humans' visual attention, can be considered as a binary classification problem between center and surround classes. Discriminant power of features for the classification is measured as mutual…

Computer Vision and Pattern Recognition · Computer Science 2013-01-18 Anh Cat Le Ngo , Kenneth Ang Li-Minn , Guoping Qiu , Jasmine Seng Kah-Phooi

A discrete system's heterogeneity is measured by the R\'enyi heterogeneity family of indices (also known as Hill numbers or Hannah--Kay indices), whose units are {the numbers equivalent}. Unfortunately, numbers equivalent heterogeneity…

Machine Learning · Statistics 2020-04-07 Abraham Nunes , Martin Alda , Timothy Bardouille , Thomas Trappenberg

Multivariate measurements taken at different spatial locations occur frequently in practice. Proper analysis of such data needs to consider not only dependencies on-sight but also dependencies in and in-between variables as a function of…

Methodology · Statistics 2024-04-12 Christoph Muehlmann , Peter Filzmoser , Klaus Nordhausen

Marginal likelihood, also known as model evidence, is a fundamental quantity in Bayesian statistics. It is used for model selection using Bayes factors or for empirical Bayes tuning of prior hyper-parameters. Yet, the calculation of…

Methodology · Statistics 2024-09-04 Anindya Bhadra , Ksheera Sagar , David Rowe , Sayantan Banerjee , Jyotishka Datta

The input data features set for many data driven tasks is high-dimensional while the intrinsic dimension of the data is low. Data analysis methods aim to uncover the underlying low dimensional structure imposed by the low dimensional hidden…

Machine Learning · Computer Science 2019-01-30 Moshe Salhov , Ofir Lindenbaum , Yariv Aizenbud , Avi Silberschatz , Yoel Shkolnisky , Amir Averbuch

We explore the design of visualizations for values spanning multiple orders of magnitude; we call them Orders of Magnitude Values (OMVs). Visualization researchers have shown that separating OMVs into two components, the mantissa and the…

Human-Computer Interaction · Computer Science 2025-03-11 Katerina Batziakoudi , Florent Cabric , Stéphanie Rey , Jean-Daniel Fekete

The paper presents new metrics to quantify and test for (i) the equality of distributions and (ii) the independence between two high-dimensional random vectors. We show that the energy distance based on the usual Euclidean distance cannot…

Methodology · Statistics 2019-10-01 Shubhadeep Chakraborty , Xianyang Zhang

Diffusion models have achieved great success in generating high-dimensional samples across various applications. While the theoretical guarantees for continuous-state diffusion models have been extensively studied, the convergence analysis…

Machine Learning · Computer Science 2025-04-15 Zikun Zhang , Zixiang Chen , Quanquan Gu

In Data Science, entities are typically represented by single valued measurements. Symbolic Data Analysis extends this framework to more complex structures, such as intervals and histograms, that express internal variability. We propose an…

Machine Learning · Statistics 2025-12-16 Diogo Pinheiro , M. Rosário Oliveira , Igor Kravchenko , Lina Oliveira

The Kullback-Leibler (KL) divergence is a foundational measure for comparing probability distributions. Yet in multivariate settings, its single value often obscures the underlying reasons for divergence, conflating mismatches in individual…

Other Computer Science · Computer Science 2025-05-06 William Cook

Finite mixture models that allow for a broad range of potentially non-elliptical cluster distributions is an emerging methodological field. Such methods allow for the shape of the clusters to match the natural heterogeneity of the data,…

The Kullback-Leibler divergence or relative entropy is an information-theoretic measure between statistical models that play an important role in measuring a distance between random variables. In the study of complex systems, random fields…

Information Theory · Computer Science 2022-03-25 Alexandre L. M. Levada

Biological systems often exhibit a heterogeneous arrangement of objects, such as assorted nuclear chromatin patterns in a tumor, assorted species of bacteria in biofilms, or assorted aggregates of subcellular particles. Principle Component…

Quantitative Methods · Quantitative Biology 2017-08-29 David H Nguyen

Quantifying the distance between datasets is a fundamental question in mathematics and machine learning. We propose \textit{magnitude distance}, a novel distance metric defined on finite datasets using the notion of the \emph{magnitude} of…

Machine Learning · Computer Science 2026-02-10 Sahel Torkamani , Henry Gouk , Rik Sarkar

In various applications, we deal with high-dimensional positive-valued data that often exhibits sparsity. This paper develops a new class of continuous global-local shrinkage priors tailored to analyzing gamma-distributed observations where…

Methodology · Statistics 2023-11-08 Yasuyuki Hamura , Takahiro Onizuka , Shintaro Hashimoto , Shonosuke Sugasawa

We initiate a study of the following problem: Given a continuous domain $\Omega$ along with its convex hull $\mathcal{K}$, a point $A \in \mathcal{K}$ and a prior measure $\mu$ on $\Omega$, find the probability density over $\Omega$ whose…

Data Structures and Algorithms · Computer Science 2020-04-17 Jonathan Leake , Nisheeth K. Vishnoi