English
Related papers

Related papers: Topology-based goodness-of-fit tests for sliced sp…

200 papers

We develop goodness-of-fit tests for max-stable random fields, which are used to model heavy-tailed spatial data. The test statistics are constructed based on the Fourier transforms of the indicators of extreme values in the heavy-tailed…

Methodology · Statistics 2025-12-09 Ying Niu , Zhao Chen , Christina Dan Wang , Yuwei Zhao

Data quality is crucial for the successful training, generalization and performance of machine learning models. We propose to measure the quality of a subset concerning the dataset it represents, using topological data analysis techniques.…

Algebraic Topology · Mathematics 2024-10-01 Álvaro Torras-Casas , Eduardo Paluzo-Hidalgo , Rocio Gonzalez-Diaz

We present new families of goodness-of-fit tests of uniformity on a full-dimensional set $W\subset\R^d$ based on statistics related to edge lengths of random geometric graphs. Asymptotic normality of these statistics is proven under the…

Statistics Theory · Mathematics 2020-07-20 Bruno Ebner , Franz Nestmann , Matthias Schulte

Statistical analysis on object data presents many challenges. Basic summaries such as means and variances are difficult to compute. We apply ideas from topology to study object data. We present a framework for using persistence landscapes…

Methodology · Statistics 2019-12-12 Vic Patrangenaru , Peter Bubenik , Robert L. Paige , Daniel Osborne

In topological data analysis, we want to discern topological and geometric structure of data, and to understand whether or not certain features of data are significant as opposed to simply random noise. While progress has been made on…

Computational Geometry · Computer Science 2020-01-10 So Mang Han , Taylor Okonek , Nikesh Yadav , Xiaojun Zheng

We propose a family of tests to assess the goodness-of-fit of a high-dimensional generalized linear model. Our framework is flexible and may be used to construct an omnibus test or directed against testing specific non-linearities and…

Methodology · Statistics 2019-11-14 Jana Janková , Rajen D. Shah , Peter Bühlmann , Richard J. Samworth

We derived a geometric goodness-of-fit index, similar to $R^2$ using topological data analysis techniques. We build the Vietoris-Rips complex from the data-cloud projected onto each variable. Estimating the area of the complex and their…

Computation · Statistics 2020-12-11 Alberto Hernández , Maikol Solís , Ronald Zúñiga

We introduce a new goodness-of-fit test for regular vine (R-vine) copula models. R-vine copulas are a very flexible class of multivariate copulas based on a pair-copula construction (PCC). The test arises from the information matrix…

Computation · Statistics 2013-06-05 Ulf Schepsmeier

We introduce a general framework for testing goodness-of-fit for Gaussian graphical models in both the low- and high-dimensional settings. This framework is based on a novel algorithm for generating exchangeable copies by conditioning on…

Methodology · Statistics 2025-01-07 Xiaotong Lin , Weihao Li , Fangqiao Tian , Dongming Huang

Classical tests of goodness-of-fit aim to validate the conformity of a postulated model to the data under study. Given their inferential nature, they can be considered a crucial step in confirmatory data analysis. In their standard…

Methodology · Statistics 2022-04-06 Sara Algeri , Xiangyu Zhang

In many applications, we encounter data on Riemannian manifolds such as torus and rotation groups. Standard statistical procedures for multivariate data are not applicable to such data. In this study, we develop goodness-of-fit testing and…

Methodology · Statistics 2021-03-02 Wenkai Xu , Takeru Matsuda

A variety of statistics based on sample spacings has been studied in the literature for testing goodness-of-fit to parametric distributions. To test the goodness-of-fit to a nonparametric class of univariate shape-constrained densities,…

Statistics Theory · Mathematics 2024-10-28 Kwun Chuen Gary Chan , Hok Kan Ling , Chuan-Fa Tang , Sheung Chi Phillip Yam

Machine learning models for repeated measurements are limited. Using topological data analysis (TDA), we present a classifier for repeated measurements which samples from the data space and builds a network graph based on the data topology.…

Machine Learning · Computer Science 2019-04-08 Henri Riihimäki , Wojciech Chachólski , Jakob Theorell , Jan Hillert , Ryan Ramanujam

Persistent homology analysis provides means to capture the connectivity structure of data sets in various dimensions. On the mathematical level, by defining a metric between the objects that persistence attaches to data sets, we can…

Machine Learning · Computer Science 2019-06-12 Henri Riihimäki , José Licón-Saláiz

In this review, the state-of-the-art for goodness-of-fit testing for spatial point processes is summarized. Test statistics based on classical functional summary statistics and recent contributions from topological data analysis are…

Methodology · Statistics 2025-01-08 Chiara Fend , Claudia Redenbach

This paper proposes a novel two-step strategy for testing the goodness-of-fit of parametric regression models in ultra-high dimensional sparse settings, where the predictor dimension far exceeds the sample size. This regime usually renders…

Methodology · Statistics 2025-12-30 Falong Tan , Jie Liu , Heng Peng , Lixing Zhu

Using the fact that some depth functions characterize certain family of distribution functions, and under some mild conditions, distribution of the depth is continuous, we have constructed several new multivariate goodness of fit tests…

Statistics Theory · Mathematics 2024-05-14 Rahul Singh , Subhajit Dutta , Neeraj Misra

We study distributions of persistent homology barcodes associated to taking subsamples of a fixed size from metric measure spaces. We show that such distributions provide robust invariants of metric measure spaces, and illustrate their use…

Computational Geometry · Computer Science 2014-01-20 Andrew J. Blumberg , Itamar Gal , Michael A. Mandell , Matthew Pancia

Modern representation learning increasingly relies on unsupervised and self-supervised methods trained on large-scale unlabeled data. While these approaches achieve impressive generalization across tasks and domains, evaluating embedding…

Vine copulas are a useful statistical tool to describe the dependence structure between several random variables, especially when the number of variables is very large. When modeling data with vine copulas, one often is confronted with a…

Methodology · Statistics 2017-05-10 Matthias Killiches , Daniel Kraus , Claudia Czado