English
Related papers

Related papers: Tree-Wasserstein Distance for High Dimensional Dat…

200 papers

We present a framework that allows for the non-asymptotic study of the $2$-Wasserstein distance between the invariant distribution of an ergodic stochastic differential equation and the distribution of its numerical approximation in the…

Machine Learning · Statistics 2021-09-27 J. M. Sanz-Serna , Konstantinos C. Zygalakis

Clustering is a data analysis method for extracting knowledge by discovering groups of data called clusters. Among these methods, state-of-the-art density-based clustering methods have proven to be effective for arbitrary-shaped clusters.…

Machine Learning · Computer Science 2023-10-26 Nabil El Malki , Robin Cugny , Olivier Teste , Franck Ravat

Sliced Wasserstein distances preserve properties of classic Wasserstein distances while being more scalable for computation and estimation in high dimensions. The goal of this work is to quantify this scalability from three key aspects: (i)…

Machine Learning · Statistics 2022-10-18 Sloan Nietert , Ritwik Sadhu , Ziv Goldfeld , Kengo Kato

Despite the notable advancements in numerous Transformer-based models, the task of long multi-horizon time series forecasting remains a persistent challenge, especially towards explainability. Focusing on commonly used saliency maps in…

Machine Learning · Computer Science 2023-09-18 Nghia Duong-Trung , Duc-Manh Nguyen , Danh Le-Phuoc

How to extract useful insights from data is always a challenge, especially if the data is multidimensional. Often, the data can be organized according to certain hierarchical structure that are stemmed either from data collection process or…

Applications · Statistics 2016-04-21 Kun Yang , Wing Hung Wong

Smoothing and filtering two-dimensional sequences are fundamental tasks in fields such as computer vision. Conventional filtering algorithms often rely on the selection of the filtering window, limiting their applicability in certain…

Information Theory · Computer Science 2025-07-22 Xufeng Chen , Liang Yan , Xiaoshan Gao

Lightweight Temporal Compression (LTC) is among the lossy stream compression methods that provide the highest compression rate for the lowest CPU and memory consumption. As such, it is well suited to compress data streams in…

Information Theory · Computer Science 2018-11-27 Bo Li , Omid Sarbishei , Hosein Nourani , Tristan Glatard

While time series prediction is an important, actively studied problem, the predictive accuracy of time series models is complicated by non-stationarity. We develop a fast and effective approach to allow for non-stationarity in the…

Applications · Statistics 2015-12-10 Daniel M. McCarthy , Shane T. Jensen

Trajectory similarity is a cornerstone of trajectory data management and analysis. Traditional similarity functions often suffer from high computational complexity and a reliance on specific distance metrics, prompting a shift towards deep…

Databases · Computer Science 2025-04-16 Jianing Si , Haitao Yuan , Nan Jiang , Minxiao Chen , Xiao Ma , Shangguang Wang

The Gromov--Wasserstein (GW) distance and its fused extension (FGW) are powerful tools for comparing heterogeneous data. Their computation is, however, challenging since both distances are based on non-convex, quadratic optimal transport…

Machine Learning · Computer Science 2025-11-14 Moritz Piening , Robert Beinert

For more than three decades, measurement of terrace width distributions (TWDs) of vicinal crystal surfaces have been recognized as arguably the best way to determine the dimensionless strength $\tilde{A}$ of the elastic repulsion between…

Materials Science · Physics 2015-06-24 T. L. Einstein , Howard L. Richards , Saul D. Cohen , O. Pierre-Louis

Comparing probability distributions is at the crux of many machine learning algorithms. Maximum Mean Discrepancies (MMD) and Wasserstein distances are two classes of distances between probability distributions that have attracted abundant…

Machine Learning · Statistics 2023-06-01 Titouan Vayer , Rémi Gribonval

Statistical models often include thousands of parameters. However, large models decrease the investigator's ability to interpret and communicate the estimated parameters. Reducing the dimensionality of the parameter space in the estimation…

Methodology · Statistics 2022-05-16 Eric Dunipace , Lorenzo Trippa

The synergy between extremely large-scale antenna arrays and terahertz technology in sixth-generation networks establishes a near-field wideband transmission environment, enabling the generation of highly focused beams. To leverage this…

Signal Processing · Electrical Eng. & Systems 2026-03-18 Qianyu Yang , Haiyang Zhang , Francesco Guidi , Anna Guerra , Davide Dardari , Baoyun Wang , Yonina C. Eldar

Dimension reduction (DR) methods provide systematic approaches for analyzing high-dimensional data. A key requirement for DR is to incorporate global dependencies among original and embedded samples while preserving clusters in the…

Machine Learning · Statistics 2023-03-10 Antoine Collas , Titouan Vayer , Rémi Flamary , Arnaud Breloy

In recent years, reversible data hiding has attracted much more attention than before. Reversibility signifies that the original media can be recovered without any loss from the marked media after extracting the embedded message. This paper…

Cryptography and Security · Computer Science 2011-02-09 Xu-Ren Luo , Chen-Hui Jerry Lin , Te-Lung Yin

Determining the appropriate number of clusters in unsupervised learning is a central problem in statistics and data science. Traditional validity indices such as Calinski-Harabasz, Silhouette, and Davies-Bouldin-depend on centroid-based…

Machine Learning · Statistics 2025-10-17 Mohammed Baragilly , Hend Gabr

Discrete diffusion models have recently shown great promise for modeling complex discrete data, with masked diffusion models (MDMs) offering a compelling trade-off between quality and generation speed. MDMs denoise by progressively…

Machine Learning · Computer Science 2026-04-15 Tianyu Xie , Shuchen Xue , Zijin Feng , Tianyang Hu , Jiacheng Sun , Zhenguo Li , Cheng Zhang

Multidimensional scaling is a statistical process that aims to embed high dimensional data into a lower-dimensional space; this process is often used for the purpose of data visualisation. Common multidimensional scaling algorithms tend to…

Machine Learning · Computer Science 2022-02-25 Pierre Lambert , Cyril de Bodt , Michel Verleysen , John Lee

Datasets with non-trivial large scale topology can be hard to embed in low-dimensional Euclidean space with existing dimensionality reduction algorithms. We propose to model topologically complex datasets using vector bundles, in such a way…

Computational Geometry · Computer Science 2023-03-17 Luis Scoccola , Jose A. Perea
‹ Prev 1 8 9 10 Next ›