English
Related papers

Related papers: Measuring Time-Series Dataset Similarity using Was…

200 papers

Generative adversarial nets (GANs) and variational auto-encoders have significantly improved our distribution modeling capabilities, showing promise for dataset augmentation, image-to-image translation and feature learning. However, to…

Allocation of personnel and material resources is highly sensible in the case of firefighter interventions. This allocation relies on simulations to experiment with various scenarios. The main objective of this allocation is the global…

Machine Learning · Computer Science 2025-07-30 Michael Corbeau , Emmanuelle Claeys , Mathieu Serrurier , Pascale Zaraté

Distances between probability distributions are a key component of many statistical machine learning tasks, from two-sample testing to generative modeling, among others. We introduce a novel distance between measures that compares them…

Machine Learning · Statistics 2025-07-09 Arturo Castellanos , Anna Korba , Pavlo Mozharovskyi , Hicham Janati

Distances between probability distributions that take into account the geometry of their sample space,like the Wasserstein or the Maximum Mean Discrepancy (MMD) distances have received a lot of attention in machine learning as they can, for…

Machine Learning · Computer Science 2020-04-29 Gaëtan Hadjeres , Frank Nielsen

Nonparametric two sample or homogeneity testing is a decision theoretic problem that involves identifying differences between two random variables without making parametric assumptions about their underlying distributions. The literature is…

Statistics Theory · Mathematics 2015-10-14 Aaditya Ramdas , Nicolas Garcia , Marco Cuturi

Measuring the distance between ontological elements is fundamental for ontology matching. String-based distance metrics are notorious for shallow syntactic matching. In this exploratory study, we investigate Wasserstein distance targeting…

Artificial Intelligence · Computer Science 2022-09-22 Yuan An , Alex Kalinowski , Jane Greenberg

This paper proposes a novel similarity measure for clustering sequential data. We first construct a common state-space by training a single probabilistic model with all the sequences in order to get a unified representation for the dataset.…

Machine Learning · Computer Science 2010-04-13 Darío García-García , Emilio Parrado-Hernández , Fernando Díaz-de-María

We introduce a new approach to nonlinear sufficient dimension reduction in cases where both the predictor and the response are distributional data, modeled as members of a metric space. Our key step is to build universal kernels…

Methodology · Statistics 2023-04-26 Qi Zhang , Bing Li , Lingzhou Xue

In object detection, a well-defined similarity metric can significantly enhance model performance. Currently, the IoU-based similarity metric is the most commonly preferred choice for detectors. However, detectors using IoU as a similarity…

Computer Vision and Pattern Recognition · Computer Science 2025-11-03 Ziqian Guan , Xieyi Fu , Pengjun Huang , Hengyuan Zhang , Hubin Du , Yongtao Liu , Yinglin Wang , Qang Ma

High-dimensional data often exhibit hierarchical structures in both modes: samples and features. Yet, most existing approaches for hierarchical representation learning consider only one mode at a time. In this work, we propose an…

Machine Learning · Computer Science 2025-10-23 Ya-Wei Eileen Lin , Ronald R. Coifman , Gal Mishne , Ronen Talmon

Using statistical learning methods to analyze stochastic simulation outputs can significantly enhance decision-making by uncovering relationships between different simulated systems and between a system's inputs and outputs. We focus on…

Methodology · Statistics 2026-05-28 Mohammadmahdi Ghasemloo , David J. Eckman

Recently, a Wasserstein-type distance for Gaussian mixture models has been proposed. However, that framework can only be generalized to identifiable mixtures of general elliptically contoured distributions whose components come from the…

Optimization and Control · Mathematics 2025-03-19 Keyu Chen , Zetian Wang , Yunxin Zhang

In this paper we investigate the sensitivity of the LWR model on network to its parameters and to the network itself. The quantification of sensitivity is obtained by measuring the Wasserstein distance between two LWR solutions…

Numerical Analysis · Mathematics 2018-04-13 Maya Briani , Emiliano Cristiani , Elisa Iacomini

Time series classification is an increasing research topic due to the vast amount of time series data that are being created over a wide variety of fields. The particularity of the data makes it a challenging task and different approaches…

Machine Learning · Statistics 2018-06-13 Amaia Abanda , Usue Mori , Jose A. Lozano

The Wasserstein distance is a metric for assessing distributional differences. The measure originates in optimal transport theory and can be interpreted as the minimal cost of transforming one distribution into another. In this paper, the…

Applications · Statistics 2026-04-21 Markus Sauerberg

Given a pair of multivariate time-series data of the same length and dimensions, an approach is proposed to select variables and time intervals where the two series are significantly different. In applications where one time series is an…

Methodology · Statistics 2024-12-11 Kensuke Mitsuzawa , Margherita Grossi , Stefano Bortoli , Motonobu Kanagawa

Functional time series (FTS) extend traditional methodologies to accommodate data observed as functions/curves. A significant challenge in FTS consists of accurately capturing the time-dependence structure, especially with the presence of…

Statistics Theory · Mathematics 2025-04-10 Jan Nino G. Tinio , Mokhtar Z. Alaya , Salim Bouzebda

Optimal transport and Wasserstein distances are flourishing in many scientific fields as a means for comparing and connecting random structures. Here we pioneer the use of an optimal transport distance between L\'{e}vy measures to solve a…

Statistics Theory · Mathematics 2023-09-18 Marta Catalano , Hugo Lavenant , Antonio Lijoi , Igor Prünster

Clustering is an important exploratory data analysis technique to group objects based on their similarity. The widely used $K$-means clustering method relies on some notion of distance to partition data into a fewer number of groups. In the…

Machine Learning · Statistics 2022-10-14 Yubo Zhuang , Xiaohui Chen , Yun Yang

Anomalies (unusual patterns) in time-series data give essential, and often actionable information in critical situations. Examples can be found in such fields as healthcare, intrusion detection, finance, security and flight safety. In this…

Applications · Statistics 2016-08-17 Evgeny Burnaev , Vladislav Ishimtsev
‹ Prev 1 8 9 10 Next ›