English
Related papers

Related papers: Distance Assessment and Hypothesis Testing of High…

200 papers

Permutation tests are a distribution free way of performing hypothesis tests. These tests rely on the condition that the observed data are exchangeable among the groups being tested under the null hypothesis. This assumption is easily…

Methodology · Statistics 2017-12-14 Daniell Toth

We propose a new nonparametric test for the supposition of independence between two continuous random variables. The test is based on the size of the longest increasing subsequence of a random permutation. We identified the independence…

Methodology · Statistics 2015-03-13 Jesus E. Garcia , Veronica A. Gonzalez-Lopez

Using variational autoencoders trained on known physics processes, we develop a one-sided threshold test to isolate previously unseen processes as outlier events. Since the autoencoder training does not depend on any specific new physics…

High Energy Physics - Experiment · Physics 2019-06-14 Olmo Cerri , Thong Q. Nguyen , Maurizio Pierini , Maria Spiropulu , Jean-Roch Vlimant

Electron and scanning probe microscopy produce vast amounts of data in the form of images or hyperspectral data, such as EELS or 4D STEM, that contain information on a wide range of structural, physical, and chemical properties of…

Machine Learning · Computer Science 2023-06-09 Arpan Biswas , Maxim Ziatdinov , Sergei V. Kalinin

Consider a set of multiple, multimodal sensors capturing a complex system or a physical phenomenon of interest. Our primary goal is to distinguish the underlying sources of variability manifested in the measured data. The first step in our…

Computer Vision and Pattern Recognition · Computer Science 2016-11-28 Vardan Papyan , Ronen Talmon

Leveraging the Wasserstein distance -- a summation of sample-wise transport distances in data space -- is advantageous in many applications for measuring support differences between two underlying density functions. However, when supports…

Machine Learning · Computer Science 2025-11-18 Cheongjae Jang , Jonghyun Won , Soyeon Jun , Chun Kee Chung , Keehyoung Joo , Yung-Kyun Noh

We propose using a permutation test to detect discontinuities in an underlying economic model at a known cutoff point. Relative to the existing literature, we show that this test is well suited for event studies based on time-series data.…

Econometrics · Economics 2022-07-12 Federico A. Bugni , Jia Li , Qiyuan Li

The assumption of conditional independence among observed variables, primarily used in the Variational Autoencoder (VAE) decoder modeling, has limitations when dealing with high-dimensional datasets or complex correlation structures among…

Machine Learning · Computer Science 2023-10-26 Seunghwan An , Jong-June Jeon

In settings requiring synthetic data generation based on a clinical cohort, e.g., due to data protection regulations, heterogeneity across individuals might be a nuisance that we need to control or faithfully preserve. The sources of such…

Machine Learning · Computer Science 2023-12-14 Kiana Farhadyar , Federico Bonofiglio , Maren Hackenberg , Daniela Zoeller , Harald Binder

The concepts of similarity and distance are crucial in data mining. We consider the problem of defining the distance between two data sets by comparing summary statistics computed from the data sets. The initial definition of our distance…

Data Structures and Algorithms · Computer Science 2019-02-05 Nikolaj Tatti

Deep Learning performs well when training data densely covers the experience space. For complex problems this makes data collection prohibitively expensive. We propose to intelligently select samples when constructing data sets in order to…

Computer Vision and Pattern Recognition · Computer Science 2020-04-01 Mark Philip Philipsen , Thomas Baltzer Moeslund

Consider an unlimited homogeneous medium disturbed by points generated via Poisson process. The neighborhood of a point plays an important role in spatial statistics problems. Here, we obtain analytically the distance statistics to $k$th…

Statistical Mechanics · Physics 2015-08-11 Cristiano Roberto Fabri Granzotti , Alexandre Souto Martinez

Machine learning models provide statistically impressive results which might be individually unreliable. To provide reliability, we propose an Epistemic Classifier (EC) that can provide justification of its belief using support from the…

Machine Learning · Computer Science 2020-10-20 Chitresh Bhushan , Zhaoyuan Yang , Nurali Virani , Naresh Iyer

One of the fundamental problems in machine learning is the estimation of a probability distribution from data. Many techniques have been proposed to study the structure of data, most often building around the assumption that observations…

Machine Learning · Statistics 2013-02-22 Oren Rippel , Ryan Prescott Adams

Noting the importance of the latent variables in inference and learning, we propose a novel framework for autoencoders based on the homeomorphic transformation of latent variables, which could reduce the distance between vectors in the…

Machine Learning · Computer Science 2019-06-04 Jaehoon Cha , Kyeong Soo Kim , Sanghyuk Lee

Synthetic data generation is of great interest in diverse applications, such as for privacy protection. Deep generative models, such as variational autoencoders (VAEs), are a popular approach for creating such synthetic datasets from…

Machine Learning · Statistics 2021-05-17 Kiana Farhadyar , Federico Bonofiglio , Daniela Zoeller , Harald Binder

Testing for the equality of two high-dimensional distributions is a challenging problem, and this becomes even more challenging when the sample size is small. Over the last few decades, several graph-based two-sample tests have been…

Methodology · Statistics 2019-11-22 Soham Sarkar , Rahul Biswas , Anil K. Ghosh

We develop a general framework for statistical inference with the 1-Wasserstein distance. Recently, the Wasserstein distance has attracted considerable attention and has been widely applied to various machine learning tasks because of its…

Statistics Theory · Mathematics 2022-02-16 Masaaki Imaizumi , Hirofumi Ota , Takuo Hamaguchi

Influence diagnosis is important since presence of influential observations could lead to distorted analysis and misleading interpretations. For high-dimensional data, it is particularly so, as the increased dimensionality and complexity…

Statistics Theory · Mathematics 2013-11-27 Junlong Zhao , Chenlei Leng , Lexin Li , Hansheng Wang

Maximum Mean Discrepancy (MMD) has been widely used in the areas of machine learning and statistics to quantify the distance between two distributions in the $p$-dimensional Euclidean space. The asymptotic property of the sample MMD has…

Statistics Theory · Mathematics 2023-08-29 Hanjia Gao , Xiaofeng Shao
‹ Prev 1 3 4 5 6 7 10 Next ›