English
Related papers

Related papers: Matrix dissimilarities based on differences in mom…

200 papers

Temporal networks are commonly used to represent dynamical complex systems like social networks, simultaneous firing of neurons, human mobility or public transportation. Their dynamics may evolve on multiple time scales characterising for…

Physics and Society · Physics 2024-02-27 Elsa Andres , Alain Barrat , Márton Karsai

The comparison of local characteristics of two random processes can shed light on periods of time or space at which the processes differ the most. This paper proposes a method that learns about regions with a certain volume, where the…

Methodology · Statistics 2022-09-14 Miguel de Carvalho , Gabriel Martos Venturini

This paper introduces a new framework to quantify distance between finite sets with uncertainty present, where probability distributions determine the locations of individual elements. Combining this with a Bayesian change point detection…

Statistical Finance · Quantitative Finance 2021-12-28 Nick James , Max Menzies

We propose a new class of metrics on sets, vectors, and functions that can be used in various stages of data mining, including exploratory data analysis, learning, and result interpretation. These new distance functions unify and generalize…

Classification in the dissimilarity space has become a very active research area since it provides a possibility to learn from data given in the form of pairwise non-metric dissimilarities, which otherwise would be difficult to cope with.…

We propose a novel method to determine the dissimilarity between subjects for functional data clustering. Spline smoothing or interpolation is common to deal with data of such type. Instead of estimating the best-representing curve for each…

Methodology · Statistics 2021-03-23 ShengLi Tzeng , Christian Hennig , Yu-Fen Li , Chien-Ju Lin

The concepts of similarity and distance are crucial in data mining. We consider the problem of defining the distance between two data sets by comparing summary statistics computed from the data sets. The initial definition of our distance…

Data Structures and Algorithms · Computer Science 2019-02-05 Nikolaj Tatti

Classification with a sparsity constraint on the solution plays a central role in many high dimensional machine learning applications. In some cases, the features can be grouped together so that entire subsets of features can be selected or…

Machine Learning · Computer Science 2014-09-05 Nikhil Rao , Robert Nowak , Christopher Cox , Timothy Rogers

In many applications, different populations are compared using data that are sampled in a biased manner. Under sampling biases, standard methods that estimate the difference between the population means yield unreliable inferences. Here we…

Statistics Theory · Mathematics 2019-11-12 Dave Zachariah , Petre Stoica

This work briefly explores the possibility of approximating spatial distance (alternatively, similarity) between data points using the Isolation Forest method envisioned for outlier detection. The logic is similar to that of isolation: the…

Machine Learning · Statistics 2019-11-26 David Cortes

Behavioral science researchers have shown strong interest in disaggregating within-person relations from between-person differences (stable traits) using longitudinal data. In this paper, we propose a method of within-person variability…

Methodology · Statistics 2025-01-08 Satoshi Usami

The Jaccard similarity index is an important measure of the overlap of two sets, widely used in machine learning, computational genomics, information retrieval, and many other areas. We design and implement SimilarityAtScale, the first…

Computational Engineering, Finance, and Science · Computer Science 2020-11-12 Maciej Besta , Raghavendra Kanakagiri , Harun Mustafa , Mikhail Karasikov , Gunnar Rätsch , Torsten Hoefler , Edgar Solomonik

Cells can utilize chemical communication to exchange information and coordinate their behavior in the presence of noise. Communication can reduce noise to shape a collective response, or amplify noise to generate distinct phenotypic…

Molecular Networks · Quantitative Biology 2019-09-24 David T. Gonzales , T-Y Dora Tang , Christoph Zechner

Diffusion models are one of the key architectures of generative AI. Their main drawback, however, is the computational costs. This study indicates that the concept of sparsity, well known especially in statistics, can provide a pathway to…

Machine Learning · Computer Science 2025-09-26 Mahsa Taheri , Johannes Lederer

Uncovering how inequality emerges from human interaction is imperative for just societies. Here we show that the way social groups interact in face-to-face situations can enable the emergence of disparities in the visibility of social…

Physics and Society · Physics 2022-03-17 Marcos Oliveira , Fariba Karimi , Maria Zens , Johann Schaible , Mathieu Génois , Markus Strohmaier

Matching is an important tool in causal inference. The method provides a conceptually straightforward way to make groups of units comparable on observed characteristics. The use of the method is, however, limited to situations where the…

Methodology · Statistics 2019-06-18 Fredrik Sävje , Michael J. Higgins , Jasjeet S. Sekhon

We introduce genetic algorithms as a means to estimate the accuracy required to discriminate among different models using experimental observables. We exemplify the technique in the context of the minimal supersymmetric standard model. If…

High Energy Physics - Phenomenology · Physics 2009-11-10 B. C. Allanach , D. Grellscheid , F. Quevedo

Quantifying the dissimilarity of two texts is an important aspect of a number of natural language processing tasks, including semantic information retrieval, topic classification, and document clustering. In this paper, we compared the…

Computation and Language · Computer Science 2023-05-05 Benjamin Shade , Eduardo G. Altmann

Self-similarity was recently introduced as a measure of inter-class congruence for classification of actions. Herein, we investigate the dual problem of intra-class dissimilarity for classification of action styles. We introduce…

Computer Vision and Pattern Recognition · Computer Science 2017-05-23 Yuping Shen , Hassan Foroosh

Uniform bin width histograms are widely used so this data graphic should represent data as correctly as possible. Method of moments based on familiar mean, variance and Fisher-Pearson skewness cure this problem.

Methodology · Statistics 2019-09-11 James S. Weber , Nicole A. Lazar