中文
相关论文

相关论文: Sublinear Time Algorithms for Earth Mover's Distan…

200 篇论文

Sampling algorithms play a pivotal role in probabilistic AI. However, verifying if a sampler program indeed samples from the claimed distribution is a notoriously hard problem. Provably correct testers like Barbarik, Teq, Flash, CubeProbe…

数据结构与算法 · 计算机科学 2025-12-09 Rishiraj Bhattacharyya , Sourav Chakraborty , Yash Pote , Uddalok Sarkar , Sayantan Sen

Optimal transport (OT) distances between probability distributions are parameterized by the ground metric they use between observations. Their relevance for real-life applications strongly hinges on whether that ground metric parameter is…

机器学习 · 统计学 2020-11-06 Matthieu Heitz , Nicolas Bonneel , David Coeurjolly , Marco Cuturi , Gabriel Peyré

In modern relational machine learning it is common to encounter large graphs that arise via interactions or similarities between observations in many domains. Further, in many cases the target entities for analysis are actually signals on…

Earth System Models (ESMs) are the primary tools for investigating future Earth system states at time scales from decades to centuries, especially in response to anthropogenic greenhouse gas release. State-of-the-art ESMs can reproduce the…

机器学习 · 计算机科学 2023-06-05 Maximilian Gelbrecht , Alistair White , Sebastian Bathiany , Niklas Boers

The earth mover's distance (EMD), also known as the 1-Wasserstein metric, measures the minimum amount of work required to transform one probability distribution into another. The EMD can be naturally generalized to measure the "distance"…

统计理论 · 数学 2024-12-11 William Q. Erickson

We define two minimum distance estimators for dependent data by minimizing some approximated Maximum Mean Discrepancy distances between the true empirical distribution of observations and their assumed (parametric) model distribution. When…

统计方法学 · 统计学 2026-01-19 Pierre Alquier , Jean-David Fermanian , Benjamin Poignard

Although ubiquitous in the sciences, histogram data have not received much attention by the Deep Learning community. Whilst regression and classification tasks for scalar and vector data are routinely solved by neural networks, a principled…

机器学习 · 计算机科学 2021-06-07 Florian List

Geometric data sets arising in modern applications are often very large and change dynamically over time. A popular framework for dealing with such data sets is the evolving data framework, where a discrete structure continuously varies…

计算几何 · 计算机科学 2025-04-28 Aditya Acharya , David M. Mount

We propose a simple subsampling scheme for fast randomized approximate computation of optimal transport distances. This scheme operates on a random subset of the full data and can use any exact algorithm as a black-box back-end, including…

统计计算 · 统计学 2020-12-17 Max Sommerfeld , Jörn Schrieber , Yoav Zemel , Axel Munk

Tracking algorithms such as the Kalman filter aim to improve inference performance by leveraging the temporal dynamics in streaming observations. However, the tracking regularizers are often based on the $\ell_p$-norm which cannot account…

信号处理 · 电气工程与系统科学 2020-05-20 Nicholas P. Bertrand , Adam S. Charles , John Lee , Pavel B. Dunn , Christopher J. Rozell

Machine learning systems operate under the assumption that training and test data are sampled from a fixed probability distribution. However, this assumptions is rarely verified in practice, as the conditions upon which data was acquired…

机器学习 · 计算机科学 2025-07-09 Eduardo Fernandes Montesuma , Fred Maurice Ngolè Mboula , Antoine Souloumiac

In this paper, we study estimators for geometric optimization problems in the sublinear geometric model. In this model, we have oracle access to a point set with size $n$ in a discrete space $[\Delta]^d$, where queries can be made to an…

计算几何 · 计算机科学 2025-04-23 Anne Driemel , Morteza Monemizadeh , Eunjin Oh , Frank Staals , David P. Woodruff

A fundamental assumption of most machine learning algorithms is that the training and test data are drawn from the same underlying distribution. However, this assumption is violated in almost all practical applications: machine learning…

机器学习 · 计算机科学 2021-12-02 Marvin Zhang , Henrik Marklund , Nikita Dhawan , Abhishek Gupta , Sergey Levine , Chelsea Finn

The earth mover's distance (EMD) is a well-known metric on spaces of histograms; roughly speaking, the EMD measures the minimum amount of work required to equalize two histograms. The EMD has a natural generalization that compares an…

组合数学 · 数学 2023-06-22 William Q. Erickson

A fundamental notion of distance between train and test distributions from the field of domain adaptation is discrepancy distance. While in general hard to compute, here we provide the first set of provably efficient algorithms for testing…

数据结构与算法 · 计算机科学 2024-06-14 Gautam Chandrasekaran , Adam R. Klivans , Vasilis Kontonis , Konstantinos Stavropoulos , Arsen Vasilyan

We study the sublinear multivariate mean estimation problem in $d$-dimensional Euclidean space. Specifically, we aim to find the mean $\mu$ of a ground point set $A$, which minimizes the sum of squared Euclidean distances of the points in…

数据结构与算法 · 计算机科学 2025-10-07 Beatrice Bertolotti , Matteo Russo , Chris Schwiegelshohn , Sudarshan Shyam

Geographic distribution shift arises when the distribution of locations on Earth in a training dataset is different from what is seen at inference time. Using standard empirical risk minimization (ERM) in this setting can lead to uneven…

机器学习 · 计算机科学 2026-02-10 Ruth Crasto , Esther Rolf

The Wasserstein metric or earth mover's distance (EMD) is a useful tool in statistics, machine learning and computer science with many applications to biological or medical imaging, among others. Especially in the light of increasingly…

最优化与控制 · 数学 2018-01-26 Jörn Schrieber , Dominic Schuhmacher , Carsten Gottschlich

We give a general unified method that can be used for $L_1$ {\em closeness testing} of a wide range of univariate structured distribution families. More specifically, we design a sample optimal and computationally efficient algorithm for…

数据结构与算法 · 计算机科学 2015-08-25 Ilias Diakonikolas , Daniel M. Kane , Vladimir Nikishkin

Domain adaptation algorithms are designed to minimize the misclassification risk of a discriminative model for a target domain with little training data by adapting a model from a source domain with a large amount of training data. Standard…

机器学习 · 统计学 2021-07-27 Werner Zellinger , Bernhard A Moser , Susanne Saminger-Platz